Java Production Engineer
Indexed description
Production Engineer II – Java / SQL / Automation (AWS)
Location: Alpharetta, GA — Onsite, 5 days per week
Lotus Technology Group is hiring a Production Engineer II to own and improve the reliability, performance, and daily operations of business-critical production systems.
The client is a global financial technology organization serving banks, credit unions, and merchants worldwide. Their platforms process payments and core banking transactions at very high volume, which makes production stability a direct business concern rather than an internal engineering metric.
This is an L1/L2 production support role for someone who wants to do more than work the queue. You'll handle live incidents across Java services, production databases, and AWS infrastructure — and you'll be expected to lead the support function forward: tightening process, configuring the tooling properly, and removing the recurring work that shouldn't need a human at all.
What you'll do
- Own L1 and L2 production support for Java-based services and data-driven applications — incident triage, root-cause analysis, escalation management, and remediation.
- Lead improvement of the support process itself: runbook quality, escalation paths, handoff discipline, and shift coverage standards.
- Configure and tune observability tooling — not just consume it. Building dashboards, alert rules, synthetic monitors, and instrumentation in Dynatrace, CloudWatch, and related platforms.
- Build automation that removes manual operational work: alerting, self-healing actions, deployment checks, and routine maintenance.
- Write and optimize SQL for troubleshooting, data validation, reconciliation, and performance analysis.
- Improve service reliability through proactive problem management, capacity planning, and performance tuning.
- Support release and change activities: deployment readiness, rollback planning, post-release verification.
- Participate in the on-call rotation and drive down recurring incident volume and mean time to restore (MTTR).
- Document operational procedures and lead postmortems and corrective action follow-through.
What you'll bring
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 4–5 years in L1/L2 production support, site reliability, or operations engineering for enterprise applications.
- Strong Java fundamentals, with the ability to debug application behavior from logs, stack traces, thread dumps, and runtime metrics.
- Hands-on AWS: CloudWatch, EC2, S3, RDS/Aurora, IAM, Lambda, Systems Manager, EKS/ECS, SNS/SQS.
- Demonstrated experience configuring observability platforms — Dynatrace, CloudWatch, or equivalent — including alert design, dashboard build-out, and instrumentation. Consuming existing dashboards is not enough for this role.
- Strong SQL — joins, aggregations, query optimization — and real experience with relational databases in production.
- Automation experience using scripting (Python, Bash, or PowerShell) and/or CI/CD pipelines.
- Working knowledge of incident management, problem management, change control, and post-incident reviews.
- Evidence of having improved a support process, not just operated within one.
- Clear communication under pressure and the ability to coordinate across teams during an active incident.
Nice to have
- Infrastructure as Code (Terraform, CloudFormation) and configuration management.
- API gateways, message queues and streams, and distributed system troubleshooting.
- A track record of moving operational KPIs — MTTR, change failure rate, availability, incident volume.
- Financial services or other regulated-industry experience.
What success looks like
In the first year, you'll have measurably reduced recurring incidents through automation and preventative fixes, shortened time-to-detect and time-to-recover through better-configured observability, and left the support process in better shape than you found it.
Working model and on-call
This role is onsite in Alpharetta, GA, five days per week. Participation in an on-call rotation is required, and occasional after-hours support may be needed for high-severity incidents or planned changes.
Start date
The client is looking to fill this role quickly. Candidates who can start immediately will be
prioritized.
To apply
Send your resume to [email protected] with the subject line "Production Engineer II – Application."
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search