Back to search
Lotus TG LLC Linkedin · Posted 8d ago

Java Production Engineer

Alpharetta, Georgia, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description


Production Engineer II – Java / SQL / Automation (AWS)

Location: Alpharetta, GA — Onsite, 5 days per week


Lotus Technology Group is hiring a Production Engineer II to own and improve the reliability, performance, and daily operations of business-critical production systems.


The client is a global financial technology organization serving banks, credit unions, and merchants worldwide. Their platforms process payments and core banking transactions at very high volume, which makes production stability a direct business concern rather than an internal engineering metric.

This is an L1/L2 production support role for someone who wants to do more than work the queue. You'll handle live incidents across Java services, production databases, and AWS infrastructure — and you'll be expected to lead the support function forward: tightening process, configuring the tooling properly, and removing the recurring work that shouldn't need a human at all.


What you'll do

  • Own L1 and L2 production support for Java-based services and data-driven applications — incident triage, root-cause analysis, escalation management, and remediation.
  • Lead improvement of the support process itself: runbook quality, escalation paths, handoff discipline, and shift coverage standards.
  • Configure and tune observability tooling — not just consume it. Building dashboards, alert rules, synthetic monitors, and instrumentation in Dynatrace, CloudWatch, and related platforms.
  • Build automation that removes manual operational work: alerting, self-healing actions, deployment checks, and routine maintenance.
  • Write and optimize SQL for troubleshooting, data validation, reconciliation, and performance analysis.
  • Improve service reliability through proactive problem management, capacity planning, and performance tuning.
  • Support release and change activities: deployment readiness, rollback planning, post-release verification.
  • Participate in the on-call rotation and drive down recurring incident volume and mean time to restore (MTTR).
  • Document operational procedures and lead postmortems and corrective action follow-through.

What you'll bring

  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
  • 4–5 years in L1/L2 production support, site reliability, or operations engineering for enterprise applications.
  • Strong Java fundamentals, with the ability to debug application behavior from logs, stack traces, thread dumps, and runtime metrics.
  • Hands-on AWS: CloudWatch, EC2, S3, RDS/Aurora, IAM, Lambda, Systems Manager, EKS/ECS, SNS/SQS.
  • Demonstrated experience configuring observability platforms — Dynatrace, CloudWatch, or equivalent — including alert design, dashboard build-out, and instrumentation. Consuming existing dashboards is not enough for this role.
  • Strong SQL — joins, aggregations, query optimization — and real experience with relational databases in production.
  • Automation experience using scripting (Python, Bash, or PowerShell) and/or CI/CD pipelines.
  • Working knowledge of incident management, problem management, change control, and post-incident reviews.
  • Evidence of having improved a support process, not just operated within one.
  • Clear communication under pressure and the ability to coordinate across teams during an active incident.

Nice to have

  • Infrastructure as Code (Terraform, CloudFormation) and configuration management.
  • API gateways, message queues and streams, and distributed system troubleshooting.
  • A track record of moving operational KPIs — MTTR, change failure rate, availability, incident volume.
  • Financial services or other regulated-industry experience.


What success looks like

In the first year, you'll have measurably reduced recurring incidents through automation and preventative fixes, shortened time-to-detect and time-to-recover through better-configured observability, and left the support process in better shape than you found it.


Working model and on-call

This role is onsite in Alpharetta, GA, five days per week. Participation in an on-call rotation is required, and occasional after-hours support may be needed for high-severity incidents or planned changes.

Start date

The client is looking to fill this role quickly. Candidates who can start immediately will be

prioritized.

To apply

Send your resume to [email protected] with the subject line "Production Engineer II – Application."

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search