Senior Data Engineer
Indexed description
Why This Role Exists
PubX builds next-generation Agentic Advertising infrastructure. Our Agentic AI makes real-time, revenue-critical decisions for digital publishers and advertisers. Our Bid Intelligence uses machine learning to optimize every programmatic ad auction individually, generating measurable revenue uplift for publishers. We've priced over 1 trillion programmatic auctions, we’re currently ranked #5 globally in Prebid Analytics Adapter Rankings, and growing.
The problem we’re solving
Digital publishers and advertisers leave significant revenue on the table because ad sales are still largely manual, static, or rule-based. The reality is every campaign is unique, and most ad campaign management systems lack the intelligence to evaluate context, demand, and deal potential in real time.
As a founding member of AgenticAdvertising.org, we're building the next generation of autonomous advertising infrastructure.
Who You Are
- A person with a pragmatic engineering mindset: you know when to use an agent framework, when a simpler automation workflow is better, and how to balance model capability, reliability, latency, cost, and user experience.
- Communicate technical ideas well in writing and conversation to both technical and non-technical audiences.
- Have taste for clean, well-tested code with thoughtful abstractions that’s easy to extend and operate with agents.
- Learn quickly when things are unfamiliar by prototyping, then hardening and documenting what you ship.
What You'll Work On
Tech: Python, SQL, Spark, Airflow, dbt, Kafka/Kinesis/SQS/Warpstream, AWS, Terraform/CDK, modern ETL
- Design and maintain high-volume data pipelines (batch + streaming) powering agentic AI features and core product workflows
- Build event-driven components using Kafka and message queues, including idempotency patterns, replay strategies, and backfill mechanisms
- Develop data models and transformation layers (lakehouse patterns, dbt-style modeling) supporting both analytics and ML/AI consumption
- Own data quality and reliability: schema management, validation, lineage, SLAs, and incident response
- Enable AI/ML workflows with robust datasets for training, evaluation, feature generation, and feedback loops from production agents
- Deploy and operate data infrastructure on AWS using infrastructure-as-code
What We’re Looking For
We’re looking for an experienced engineer who has worked on production systems and enjoys solving practical problems with AI.
You’ve likely have:
- Strong data engineering fundamentals: data modeling, partitioning, performance tuning, and cost-aware design for high-volume workloads
- Experience buildingstreaming and event-driven systems(Kafka/queues), including handling late/out-of-order events, backfills, and real-world data edge cases
- Strong SQL + Python skills, and comfort with modern data stack tooling (e.g., Spark, Airflow/Dagster, dbt, warehouse/lakehouse patterns)
- Hands-on AWS experience with production operations for data systems: monitoring, incident response, and security considerations (PII, access control, encryption, auditability)
- Familiarity integrating data withAI/ML and agentic systems: feature pipelines, evaluation datasets, grounding/citations inputs, and feedback capture from agent outcomes
Bonus (not required):
- Experience with AdTech or other high volume real-time systems
Who This Role Will Suit
- This role suits engineers who like a mix of autonomy and collaboration, and who are comfortable working in an environment that’s still evolving.
- We’re a distributed team with a growing engineering presence in India, so comfort with async collaboration and clear written communication is important.
- We use agentic coding tools heavily (e.g. Cursor and Claude Code) to plan, scaffold, refactor, and debug production code, while maintaining strong engineering judgment and ownership of outcomes.
Company Benefits
- Competitive salary with meaningful equity
- Fully remote, async-friendly working
- Supportive, low-ego engineering culture
- Budget for learning and professional development
Interview Process
Our process is designed to be practical and respectful.
- CV & Profile Review – Relevant experience and background
- Initial Chat (30 mins) – Motivation and role fit
- Architecture Interview (60 mins) – Architecture, design choices, and real scenarios
- Agentic Coding Exercise (60 mins) – A small task related to the role, executed by agents
If you’re interested in building and shaping real systems in a growing product company, at the forefront of AdTech innovation, we’d love to hear from you.
We will process your personal data in accordance with our Recruitment Privacy Notice: https://pubx.ai/privacy/recruitment/
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search