Data Engineer
Indexed description
We go beyond open standards-based cyber ratings. Black Kite helps organizations make smarter risk decisions, strengthen business resilience, and scale their third-party cyber risk management programs in an increasingly complex digital environment. Our work has earned consistent recognition from customers and industry analysts alike.
WHY BLACK KITE
We're a fast-moving, high-impact team solving one of the most critical challenges in cybersecurity today. If you're looking to do meaningful work alongside sharp, collaborative people — and grow your career in a space that matters — you're in the right place.
About The Role
We are looking for a Mid-Level Data Engineer to join our growing cybersecurity data platform team. In this role, you will design, build, and optimize scalable data pipelines that support cyber risk intelligence, threat monitoring, and third-party security analytics. You will work closely with cybersecurity researchers, software engineers, and data scientists to process large-scale security data from multiple structured and unstructured sources.
This role is ideal for someone who enjoys solving complex data engineering problems, building reliable ETL/ELT systems, and working in a fast-paced cybersecurity environment where data quality, scalability, and automation are critical.
Key Responsibilities
Data Pipeline Development & Maintenance:
- Design, develop, and maintain scalable ETL/ELT pipelines for processing cybersecurity and risk intelligence data.
- Build and optimize batch and near real-time data ingestion systems from APIs, web data sources, internal services, and external threat intelligence feeds.
- Develop robust data transformation and normalization workflows for large-scale datasets.
- Ensure data quality, consistency, monitoring, and reliability across pipelines and storage systems.
- Improve pipeline performance, scalability, and fault tolerance in cloud-based environments.
- Work with structured and unstructured datasets including security events, vulnerabilities, third-party risk data, and internet-scale telemetry.
- Support the architecture and maintenance of data lakes, warehouses, and distributed processing systems.
- Collaborate with DevOps and infrastructure teams to improve CI/CD workflows and data deployment processes.
- Implement logging, observability, and monitoring solutions for data systems.
- Contribute to automation initiatives that improve operational efficiency and reduce manual processes.
- Partner with cybersecurity researchers, analysts, and data scientists to enable intelligence-driven products and analytics.
- Collaborate with software engineering teams to integrate data services into production systems.
- Participate in technical discussions, architecture reviews, and engineering planning sessions.
- Document data workflows, schemas, and engineering standards clearly and effectively.
- Support troubleshooting and root-cause analysis for production data issues.
- Bachelor’s degree in Computer Engineering, Computer Science, Software Engineering, Mathematics, or a related field - or equivalent practical experience.
- 3+ years of experience in Data Engineering, Backend Engineering, or related software engineering roles.
- Strong proficiency in Python for data engineering and automation tasks.
- Experience building and maintaining ETL/ELT pipelines in production environments.
- Solid understanding of SQL and relational database design principles.
- Experience working with large-scale datasets and distributed data processing systems.
- Familiarity with cloud platforms such as AWS, GCP, or Azure.
- Experience with data orchestration and workflow management tools.
- Strong analytical and problem-solving skills.
- Good written and verbal English communication skills.
- Familiarity with PostgreSQL, ClickHouse, Elasticsearch, MongoDB, or other large-scale data storage systems.
- Understanding of data modeling, schema evolution, and data governance concepts.
- Experience working with Docker, Kubernetes, and CI/CD pipelines.
- Knowledge of cybersecurity concepts, threat intelligence, vulnerability management, or internet-scale scanning data.
- Experience processing API, streaming, or event-driven data architectures.
- Familiarity with observability and monitoring platforms.
- Experience working in Linux-based environments.
If this sounds like you, then we want to hear from you!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search