Senior Data Engineer
Indexed description
Our Team Maintains a Global-scale Geospatial Data Platform In Google BigQuery, Holding Many Terabytes Of Data Across Carbon Transition Risk, Physical Climate Risk, And Social/demographic Features — Feeding Analytical Products For Fixed Income And Real Estate Financial Instruments, Supporting The Ambitious Product Roadmap For ICE Climate And Other Data Products. Our Engineering Stack Includes
- Orchestration: Airflow, moving toward composable task abstractions over a shared pipeline framework
- Transformation: dbt, and other data lineage and DQA tools, primarily using Google BigQuery
- Geospatial processing: Python (GeoPandas, Shapely, GeoAlchemy2 against PostGIS) for vector operations, and R
- Execution and compute environments: Hybrid across Google Cloud Platform and on-premise RHEL Linux infrastructure
- Ingestion: Third-party vendor feeds via API, SFTP, cloud storage, and database replication
Responsibilities
- Take significant components of the data platform from “works” to “mature” — tightening reliability, observability, cost/performance characteristics, and operational discipline across our ingestion, transformation, and serving layers.
- Establish and foster adoption of technical standards for the team’s work — including Airflow DAG structure, dbt model layout, BigQuery schema and partitioning conventions, pipeline testing practices, and deployment workflows.
- Lead technical design discussions, mentor other data engineers through code review, pairing, and design-doc review, and grow them along their career path.
- Act as a technical point of contact for cross-functional initiatives — partnering with data science, climate science, product, and infrastructure colleagues to drive forward decisions and make tradeoffs explicit.
- Deliver day-to-day work across the stack above — authoring Airflow DAGs and dbt models, contributing geospatial processing capabilities, and shipping cleanly partitioned, audit-friendly outputs from ingestion through serving.
- Support data science and climate science teams by helping design the tooling, training, and validation environments, and by deploying their trained models into production.
- Effectively leverage AI and LLM-based developer tooling to accelerate development workflows and improve code quality.
- Identify opportunities to improve and optimize data pipelines — for speed, cost, robustness, integrity, and operational simplicity.
- Work with business analysts, product management, and adjacent engineering teams to understand and refine new data requirements.
- 5+ years of professional experience as a data engineer, with a track record of architecting, shipping, and operating production data pipelines end-to-end.
- Experience mentoring and developing other data engineers — through code review, pairing, design discussions, and career coaching.
- Ability to establish and foster adoption of technical standards.
- A habit of actively monitoring, evaluating, and prototyping emerging big-data, geospatial, and machine-learning technologies and platforms — staying conversant in advances across cloud data engines, geospatial libraries and standards, and ML/MLOps frameworks — and bringing the most promising into the team’s design discussions, evaluations, and adoption decisions.
- Strong system-design judgment across the tradeoff space of performance, cost, maintainability, and auditability
- Comfort scoping, decomposing, and delegating work for other engineers.
- Strong written and verbal communication — able to translate technical tradeoffs for senior business, product, and client stakeholders.
- Deep fluency in modern, typed Python as a primary working language, including comfort with type-driven design (e.g. Pydantic v2).
- Strong SQL background, including experience partitioning, clustering, and performance-tuning queries on modern cloud warehouses — Google BigQuery experience strongly preferred.
- Production experience with dbt for managing warehouse transformations, and with Airflow (or a comparable orchestrator) for workflow orchestration.
- Solid grounding in geospatial data engineering — Python tooling (GeoPandas, Shapely), spatial databases (PostGIS), raster processing, or adjacent skills.
- A systems-thinking orientation: anticipates cascading effects of upstream data changes, schema evolution, and vendor corrections; designs pipelines with observability, auditability, and graceful failure in mind.
- Comfort owning production incidents and debugging distributed systems.
- Experience working cooperatively with systems, network, and infrastructure engineering and operations teams to ensure proper monitoring, alerting, and incident response workflows.
- Demonstrated ability to integrate AI/LLM coding assistants productively — treating them as a force multiplier rather than a substitute for judgment.
- Curiosity about the financial and climate/geospatial domains and contexts the team operates in.
- Well-versed in and opinionated about the modern Python ecosystem.
- Exposure to columnar and lakehouse technologies (Parquet, ClickHouse, DuckDB).
- Working understanding of data lineage, data quality validation, and metadata/cataloging frameworks.
- Prior experience in a hybrid cloud + on-premise environment, and with full software development lifecycle (SDLC) best practices and processes.
- Prior exposure to ML deployment workflows — supporting data science teams with training tooling and/or model-serving infrastructure.
- Familiarity with R, particularly geospatial packages.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search