Senior Site Reliability Engineer (SRE)
Indexed description
EPAM Romania
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
We are seeking a Senior Site Reliability Engineer to build the application layer that combines single-application observability data into end-to-end insight for a real-time platform. This role applies deep technical expertise to design, build and improve the software and infrastructure that power unified observability across a global Real-Time estate.
Responsibilities
- Build and maintain application components that aggregate, correlate and present observability data from individual services into unified end-to-end views of the Real-Time platform
- Implement customer-centric aggregation and hotspot detection algorithms that reflect timeliness, completeness, accuracy and stability of market data across the Real-Time estate
- Develop and extend telemetry pipelines that ingest, transform and route metrics, traces and logs from distributed services into a coherent observability layer
- Design and build custom synthetic monitoring agents deployed across hundreds of global sites to continuously measure customer experience of Real-Time data from the edge
- Implement GitOps and API-driven workflows for observability assets to ensure consistent deployment, versioning and promotion through build pipelines
- Partner with squads to integrate observability instrumentation into applications throughout Dev, Test, PPE and Prod environments
- 5+ years of hands-on software engineering experience building data-pipeline applications such as metrics aggregation, distributed tracing or real-time streaming
- Background in latency-sensitive or market data systems
- Proficiency in at least one systems-level language (C++, Go or Rust) and one scripting/application language (Python, Java or TypeScript)
- Experience building and operating custom synthetic monitoring solutions at scale, including lightweight agents distributed across geographically diverse sites
- Familiarity with OpenTelemetry SDKs, collectors and schema conventions for instrumentation and telemetry export
- Skills in cloud-native, containerized and Kubernetes environments alongside API- and GitOps-based workflows, config-as-code and CI/CD pipelines for infrastructure
- Strong analytical mindset for modeling distributed-system behaviors combined with effective communication skills to work across squads and simplify performance concepts for stakeholders
- We believe that the greatest strength of the company is its people. EPAM is fully committed to help its employees to reach their full potential and achieve their professional goals through continues learning. With this in mind, we would like to introduce to you few of the many opportunities and services which we believe will help you expand your current knowledge:
- Full access to cutting-edge tools and technologies
- Competitive compensation depending on experience and skills
- All-around Social package: professional & soft skills training, medical & family care programs, sports
- Relocation opportunities
- Free English classes
- Unlimited access to LinkedIn learning solutions
- Continuous experience exchange with experts and professionals worldwide
- Friendly team and comfortable working environment
- Engineering, corporate, and social events within and outside the Company
- Flexible working schedule
- Opportunities for self-realization
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search