Back to search
Global Fintech Firm Linkedin · Posted 2mo ago

Site Reliability Engineer - AWS

New York City, New York, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Responsibilities

  • Automate infrastructure and operational workflows using Infrastructure as Code (IaC) with Terraform and AWS CDK.
  • Develop and optimize CI/CD pipelines using Amazon CodeBuild, GitHub Actions, and Terraform Enterprise to improve software delivery for large-scale distributed systems.
  • Implement and maintain observability solutions using Datadog and AWS CloudWatch to establish real-time monitoring and proactive incident response.
  • Ensure high availability and performance of relational and time series databases (PostgreSQL/QuestDB), including replication, failover strategies, and query optimization.
  • Deploy and maintain in-memory caching systems (Redis).
  • Deploy and maintain distributed SQL databases (i.e. CockroachDB, YugabyteDB)
  • Develop and enforce system reliability best practices by implementing SLOs/SLIs, automated recovery processes, and post-mortem analysis.
  • Leverage Apache Airflow to automate workflow orchestration and support engineering teams.
  • Manage and support distributed event streaming platforms (Apache Kafka)
  • Assist with automating integrations and workflows across Slack, Jira, and other operational tools to improve team efficiency and response times.
  • Implement security best practices for IAM, networking, and encryption to maintain a robust security posture and regulatory compliance.
  • Help with maintaining high code quality standards using tools such as Sonarqube
  • Ensure high code quality standards are maintained with the assistance of tools like Sonarqube


Qualifications

  • Extensive experience with AWS, including EC2, ECS, S3, RDS, Lambda, MWAA, API Gateway, Route53, SSM, EFS, Backup, DynamoDB, Cloudfront, VPC, CodeBuild, CodePipeline, CodeArtifact, Config, Elasticache, and networking components in a highly available environment.
  • Strong programming and automation skills in Python, including building CLI tools or automation scripts for infrastructure operations.
  • Proven experience with event streaming (i.e. Kafka/Kinesis).
  • Deep understanding of distributed systems architecture, fault tolerance, disaster recovery, and performance tuning.
  • Hands-on experience with observability and monitoring tools such as Datadog, ELK, CloudWatch.
  • Proven track record managing infrastructure with Terraform or AWS CDK, and optimizing CI/CD pipelines with Amazon CodePipeline, GitHub Actions, or Terraform Enterprise.
  • Experience managing and optimizing relational databases, including schema management and query performance tuning.
  • Strong background in incident management, system reliability, and operational leadership, with experience conducting root cause analysis and implementing reliability improvements.
  • Preferred experience in financial systems, low-latency trading, prime brokerage, or similar high-performance environments.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search