Back to search
Ahold Delhaize USA Linkedin · Posted 11d ago

Principal Platform Engineer -Infrastructure Automation & Agentic Engineering

Salisbury

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Category/Area of Expertise: IT & Technology

Job Requisition: 546158

Address: USA-NC-Salisbury-2110 Executive Drive


Ahold Delhaize USA, a division of global food retailer Ahold Delhaize, is part of the U.S. family of brands, which includes five leading omnichannel grocery brands - Food Lion, Giant Food, The GIANT Company, Hannaford and Stop & Shop. Our associates support the brands with a wide range of services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more.


Primary Purpose

The Principal Platform Engineer is an enterprise-level individual contributor within the Technology Infrastructure Hosting team. The role shapes hosting-platform vision, strategy, roadmaps, governance, priorities, investment recommendations, and delivery outcomes for complex, business-critical initiatives across multiple teams and systems. It develops reusable modules, golden paths, self-service capabilities, engineering standards, evaluation methods, and operational controls across Windows Server, Active Directory and identity integrations; AIX/POWER and Linux; VMware, Nutanix, Hyper-V and related compute; storage, backup, SAN and NAS; middleware; networking and firewalls; Datadog, Nagios and related observability; ServiceNow CMDB, ITSM and workflow integrations; and adjacent infrastructure services. The engineer remains hands-on with code and production systems and converts approved designs, domain standards, runbooks, lifecycle obligations, and recurring work into secure, scalable, reliable, cost-effective, and version-controlled capabilities using IaC, configuration management, CI/CD, deterministic orchestration, and governed AI where it adds value.


This role operates with broad decision-impacting opportunities over platform engineering priorities and outcomes without direct people-management responsibility. It aligns senior leaders and cross-functional stakeholders, establishes governance for consistent delivery, coaches senior technical leaders and engineers, resolves systemic cross-domain problems, and challenges unsupported completion claims with evidence. Engineering-whether manual, scripted, IaC-based, or AI-assisted-must improve scalability, developer and operator experience, reliability, security, lifecycle currency, recoverability, observability, auditability, service and financial outcomes.


Our flexible/hybrid work schedule includes 3 in-person days in our Salisbury, NC office and 2 remote days. Applicants must be currently authorized to work in the United States on a full-time basis.


Duties & Responsibilities

  • Platform strategy and roadmap. Define and execute the enterprise hosting-platform vision, multi-year roadmap, capability priorities, target outcomes, and investment sequencing. Use business analysis, demand, risk, lifecycle, service performance, experience, and cost data to recommend priorities and align senior stakeholders.
  • Governance and portfolio leadership. Establish decision rights, engineering standards, intake and prioritization methods, delivery controls, scorecards, and review forums for consistent platform outcomes. Lead complex initiatives across internal teams and external partners using Lean-Agile and SAFe-aligned planning, user stories, estimation, dependency management, and incremental delivery.
  • Cross-domain platform engineering. Translate approved designs into reusable implementation patterns, automation, and acceptance criteria spanning Windows/AD, AIX/Linux, VMware/Nutanix/Hyper-V, storage/backup/SAN/NAS, middleware, network/firewall, observability, and ServiceNow. Identify upstream and downstream dependencies, raise material design gaps to the accountable design authority, and verify implementation readiness.
  • Platform product and experience management. Treat shared infrastructure capabilities as products with defined users, service levels, adoption targets, roadmaps, documentation, support models, and feedback loops. Enable self-service through service catalogs, APIs, golden paths, and internal developer portal capabilities that reduce friction without weakening controls.
  • Infrastructure as Code and CI/CD. Implement approved designs with domain owners through secure IaC delivery pipelines using version control, peer review, automated validation, policy checks, plan/review/apply controls, secrets handling, artifact traceability, release gates, rollback, and drift detection. Work across Terraform or OpenTofu, Ansible, PowerShell, Python, ARM/Bicep, and applicable platform-native automation.
  • Runbook-to-automation engineering. Identify high-volume, high-risk, and toil-heavy runbooks; decompose procedures into deterministic steps; define prerequisites, approvals, validations, error handling, rollback, evidence capture, and exception paths; then deliver production-ready orchestration.
  • Generative and Agentic AI for operations. Design and implement governed AI-assisted workflows that can interpret approved runbooks, assemble execution plans, invoke tools, preserve state, request approvals, produce evidence, and stop safely when confidence, policy, or environmental conditions are not met.
  • Prompt and context engineering. Create version-controlled system instructions, task prompts, tool descriptions, retrieval/context strategies, structured outputs, and prompt test suites. Manage prompt injection, data-boundary, hallucination, and tool-misuse risks through least privilege, allowlists, approvals, and validation.
  • Agent harnesses and orchestration. Engineer the runtime scaffolding around agents, including tool interfaces, session state, memory, planning, bounded loops, approval policies, observability, error recovery, and human-in-the-loop handoffs. Separate model reasoning from deterministic control logic and privileged execution.
  • Continuous agentic improvement. Establish bounded self-improvement loops in which production telemetry, failed cases, reviewer feedback, and test results propose changes to prompts, policies, tools, or workflows. Require evaluation, versioning, peer review, approval, and controlled rollout before promotion. Do not permit unreviewed self-modification in production.
  • Evaluation and quality engineering. Build offline and pre-production evaluation harnesses for task success, tool choice, policy compliance, grounding, security, latency, cost, failure recovery, and reproducibility. Maintain representative test cases, regression suites, red-team scenarios, and release thresholds.
  • Financial and capacity stewardship. Apply financial analysis and FinOps practices to platform planning and delivery, including demand and capacity forecasting, unit-cost and consumption visibility, cost allocation, vendor and licensing trade-offs, optimization opportunities, and benefit realization. Balance resilience, performance, technical debt, experience, and cost in recommendations.
  • Security, risk, and compliance by design. Coordinate with Security and Governance to embed approved identity, secrets, least privilege, MFA, logging, audit evidence, encryption, policy-as-code, vulnerability controls, and exception-management requirements within automation and agent solutions. Align AI lifecycle controls to approved enterprise risk practices and NIST-aligned governance.
  • Reliability and observability. Define SLIs, SLOs, telemetry, traces, dashboards, alerts, run histories, change evidence, and operational health measures for pipelines and agents. Lead troubleshooting and systemic correction for high-impact platform and automation failures.
  • Platform lifecycle and resilience. Set and enforce lifecycle practices for provisioning, configuration, patching, upgrades, currency, backup, restore, high availability, disaster recovery, decommissioning, documentation, and operational reporting. Ensure lifecycle risk and recovery readiness are visible in roadmaps and governance decisions.
  • Basis Engineering and Definition of Done. Translate domain specifications into pipelines, controls, scorecards, reference implementations, and acceptance criteria. A capability is not done until authoritative inventory and ownership are recorded; supported lifecycle and compatibility are verified; security and privileged-access controls pass; testing covers normal, failure, rollback, and recovery paths; monitoring, logging, SLI/SLO and alert routing are operational; ServiceNow CI, dependency, change and knowledge records are complete; runbooks and support handoffs are current; evidence is retained; and exceptions have accountable owners and expiry dates.
  • Technical leadership and organizational capability. Set the technical bar, lead implementation and operational-readiness reviews, coach senior platform leaders and engineers, and strengthen capability across internal and supplier teams. Independently verify completion claims using artifacts, telemetry, CMDB records, test results, monitoring coverage, recovery evidence, financial outcomes, and change results; communicate material risks, trade-offs, and recommendations to senior leadership.


Salary Range: $163,280 - $244,920

Actual compensation offered to a candidate may vary based on their unique qualifications and experience, internal equity, and market conditions. Final compensation decisions will be made in accordance with company policies and applicable laws.


[Please see full job description at application link.]

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search