Back to search
Value Spectrum Technologies Linkedin · Posted 21d ago

Principle Architect || Onsite in Phoenix, AZ || W2

Phoenix, Arizona, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Position Summary

We are seeking an experienced Principal Architect to lead the design and implementation of an enterprise-grade Centralized Observability, AIOps, and Self-Healing Platform. The platform will unify metrics, logs, traces, events, and infrastructure telemetry across cloud and on-premises environments while leveraging AI/ML for anomaly detection, intelligent alerting, root cause analysis, and automated remediation.

The ideal candidate has deep expertise in Observability, Platform Engineering, Cloud Architecture, SRE, Kubernetes, Event-Driven Architecture, AI/ML integration, and Automation.

Key Responsibilities

Platform Architecture

  • Design enterprise-wide centralized observability architecture.
  • Define platform standards and reference architectures.
  • Build a multi-tenant observability platform supporting multiple customers and business units.
  • Design for high availability, scalability, and resilience.
  • Establish governance, onboarding standards, and platform lifecycle management.

Observability Architecture

Design And Standardize:

  • Metrics collection
  • Distributed tracing
  • Centralized logging
  • Event correlation
  • Synthetic monitoring
  • Real User Monitoring (RUM)
  • Infrastructure monitoring
  • Application Performance Monitoring (APM)
  • Database observability
  • Network observability

Implement Observability Using Tools Such As:

  • Prometheus
  • Grafana
  • OpenTelemetry
  • Datadog
  • Splunk
  • CloudWatch
  • Loki
  • Tempo
  • Jaeger
  • Elasticsearch

AIOps

Design AI-driven Capabilities Including:

  • Intelligent alert correlation
  • Event deduplication
  • Dynamic thresholding
  • Anomaly detection
  • Predictive analytics
  • Capacity forecasting
  • Root cause analysis
  • Incident prioritization
  • Service dependency mapping
  • Change impact analysis

Self-Healing Platform

Design Automated Remediation Workflows For:

  • Kubernetes pod failures
  • Container restarts
  • Node failures
  • Database connectivity issues
  • Memory leaks
  • Disk space issues
  • High CPU utilization
  • Service failures
  • Network issues
  • Certificate expiry
  • Auto-scaling
  • Rollback automation

Integrate With:

  • Rundeck
  • StackStorm
  • Ansible
  • Terraform
  • Kubernetes Operators
  • Argo Workflows
  • GitHub Actions
  • Jenkins

Cloud & Platform Engineering

Architect Solutions Across:

  • AWS
  • Azure
  • Google Cloud Platform
  • Kubernetes
  • OpenShift
  • Docker
  • Service Mesh (Istio, Linkerd)

Design:

  • Multi-cluster observability
  • Multi-region deployment
  • Hybrid cloud observability
  • Disaster recovery

AI Integration

Build AI Capabilities Using:

  • Large Language Models (LLMs)
  • Retrieval-Augmented Generation (RAG)
  • Vector databases
  • AI agents
  • Knowledge graphs
  • Model orchestration
  • AI-assisted runbooks
  • Automated incident summarization
  • Conversational operations assistants

Experience With:

  • OpenAI-compatible APIs
  • Amazon Bedrock
  • Azure OpenAI
  • Google Vertex AI
  • LangGraph, LangChain, or similar orchestration frameworks

Event-Driven Architecture

Design Integrations Using:

  • Apache Kafka
  • IBM MQ
  • RabbitMQ
  • Amazon EventBridge
  • Event-driven microservices

SRE Practices

Implement:

  • SLIs
  • SLOs
  • Error budgets
  • Incident management
  • Chaos engineering
  • Reliability engineering
  • Capacity planning
  • Production readiness reviews

Security

Implement:

  • RBAC
  • OAuth2 / OIDC
  • mTLS
  • Secrets management
  • Audit logging
  • Zero Trust principles
  • Compliance controls

Required Technical Skills

Observability

  • Prometheus
  • Grafana
  • OpenTelemetry
  • Datadog
  • Splunk
  • Loki
  • Tempo
  • Jaeger
  • Elasticsearch

Cloud

  • AWS
  • Azure
  • Google Cloud Platform

Kubernetes

  • Kubernetes
  • OpenShift
  • Helm
  • Argo CD

Automation

  • Terraform
  • Ansible
  • Python
  • Bash
  • PowerShell

Programming

  • Python
  • Go
  • Java
  • Node.js

AI/ML

  • LLM integration
  • RAG
  • Vector databases
  • ML-based anomaly detection
  • AI agents

Messaging

  • Kafka
  • IBM MQ
  • RabbitMQ

Databases

  • PostgreSQL
  • MySQL
  • MongoDB
  • Redis

Preferred Experience

  • Banking / Financial Services
  • Insurance
  • Healthcare
  • Enterprise SaaS
  • Multi-cloud environments
  • Large-scale production operations

Responsibilities

The successful candidate will:

  • Define the enterprise observability strategy.
  • Build a centralized telemetry platform.
  • Design AI-driven incident detection and correlation.
  • Implement automated self-healing workflows.
  • Standardize dashboards, alerts, and telemetry collection.
  • Establish SRE best practices.
  • Create reusable platform components.
  • Mentor engineering teams.
  • Lead architecture reviews.
  • Drive cloud modernization initiatives.

Nice-to-Have Skills

  • OpenTelemetry Collector customization
  • eBPF-based observability
  • FinOps
  • ServiceNow integration
  • PagerDuty integration
  • Opsgenie integration
  • Knowledge graph technologies
  • Digital twins for IT operations
  • MLOps experience

Leadership Skills

  • Enterprise architecture
  • Executive communication
  • Technical mentoring
  • Cross-functional leadership
  • Vendor evaluation
  • Strategic roadmap planning

Success Metrics

The Architect Will Be Expected To Deliver:

  • A single-pane-of-glass observability platform.
  • Reduction in Mean Time to Detect (MTTD).
  • Reduction in Mean Time to Resolve (MTTR).
  • Intelligent alert noise reduction.
  • Automated remediation for common operational issues.
  • Enterprise-wide telemetry standards.
  • High platform availability and scalability.
  • Improved developer and operations productivity.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search