Data Engineer
Indexed description
Direct End Client: California State Teachers' Retirement System (CalSTRS)
Job Title: Back-End Developer – Data Engineering, Snowflake & GenBI
Duration: 12 Months + Possible 36-Month Extension
Location: West Sacramento, CA
Work Model: Hybrid – On-site 2–3 business days per week at CalSTRS Headquarters
Hours Per Week: 40 Hours
Interview Type: Web/ In person
Ceipal ID: SCA_DEV980_PT
Job Code: RFO #5805147980
Scope of Project:
CalSTRS is seeking experienced technical resources to support the Investment Data Warehouse (IDW) platform. The project will focus on the design, development, implementation, maintenance, and support of technical solutions across the IDW platform.
The team will enhance the IDW architecture and accelerate investment-related use cases involving:
- Generative Business Intelligence (GenBI)
- Artificial Intelligence (AI)
- Machine Learning (ML)
- Agentic AI
- AWS
- Snowflake
- Enterprise data and analytics
The selected resources will collectively provide full-stack development coverage across data engineering, analytics, AI/ML, application development, APIs, semantic modeling, and cloud technologies.
A mandatory project objective is knowledge transfer to CalSTRS employees, including providing documentation, technical materials, and other requested content necessary for CalSTRS staff to maintain and operate the solutions.
Responsibilities:
1. Data Engineering & Data Warehouse
- Build, maintain, and optimize ETL/ELT pipelines.
- Design and develop cloud data warehouse solutions.
- Develop and optimize Snowflake data solutions.
- Work with Amazon Redshift and Azure Synapse or similar cloud data platforms.
- Integrate source systems and data feeds into the Investment Data Warehouse (IDW).
- Develop data warehouse objects, transformations, and integrations.
- Ensure data quality, reliability, and performance.
2. SQL & Data Modeling
- Develop advanced SQL queries and stored procedures.
- Perform advanced joins, window functions, and query performance tuning.
- Design and implement star and snowflake schemas.
- Develop enterprise data models and semantic layers.
- Support semantic models for business intelligence and AI-driven analytics.
3. ETL/ELT & Orchestration
- Develop and maintain ETL/ELT pipelines.
- Work with orchestration tools such as Apache Airflow and dbt.
- Automate data processing and transformation workflows.
- Monitor pipeline execution, failures, and performance.
- Implement appropriate logging and error-handling mechanisms.
4. GenBI, AI/ML & GenAI
- Support development of GenBI and AI-powered analytics solutions.
- Develop solutions using AI/ML and Generative AI technologies.
- Work with OpenAI APIs or comparable LLM platforms.
- Develop natural-language-to-SQL/query solutions.
- Design prompt strategies for business insights and analytics.
- Apply prompt engineering techniques to control LLM behavior.
- Support AI-assisted business intelligence and analytics use cases.
- Develop automated tests to help prevent Text-to-SQL engine hallucinations.
5. API & Application Development
- Develop APIs supporting GenBI and AI-powered solutions.
- Support backend services for analytics applications.
- Work with application teams to integrate data, AI, and analytics services.
- Support interactive analytics applications and dashboards.
- Collaborate with front-end developers working with React, Chainlit, and Streamlit.
6. Cloud & Infrastructure
- Work within AWS and Snowflake ecosystems.
- Support AWS services including Bedrock and SageMaker.
- Implement infrastructure automation using Terraform and Ansible.
- Support containerized applications using Docker and Kubernetes.
- Support CI/CD pipelines and automated deployments.
- Monitor cloud infrastructure and application performance.
7. Monitoring, Operations & Support
- Monitor data pipelines and platform performance.
- Implement monitoring and logging for ML, BI, and data systems.
- Troubleshoot data, application, and infrastructure issues.
- Support ongoing maintenance and operations of the IDW platform.
- Optimize query performance, pipeline reliability, and platform availability.
8. Collaboration & Knowledge Transfer
- Collaborate with developers, data engineers, architects, analysts, and business stakeholders.
- Participate in requirements analysis and technical solution design.
- Follow applicable SDLC and Agile Development practices.
- Adhere to CalSTRS Minimum Information Security Requirements (MISR) and AI governance processes.
- Prepare technical documentation and project materials.
- Provide knowledge transfer to CalSTRS employees.
Required/Preferred Skills:
Required Qualifications
- 7+ years of experience in the Information Technology field with extensive experience in report writing, data analysis, or database querying.
- 5+ years of experience in data engineering, data warehousing, or BI analytics.
- 2+ years of experience with ETL/ELT pipelines.
- 2+ years of experience with cloud data warehouses such as Snowflake, Amazon Redshift, Azure Synapse, and orchestration tools such as Airflow and dbt.
- 2+ years of experience with SQL, including:
- Advanced joins
- Window functions
- Performance tuning
- 2+ years of experience with data modeling, including:
- Star schemas
- Snowflake schemas
- Semantic layers
- Hands-on experience with at least one major BI tool:
- Power BI
- Tableau
- Looker
- 2+ years of hands-on experience with:
- Docker
- Kubernetes
- CI/CD pipelines
- Monitoring and logging
- Terraform
- Ansible
AI/ML & GenAI Skills
- Experience with AI/ML or Generative AI technologies.
- Experience with OpenAI APIs or similar LLM platforms preferred.
- Experience building natural-language-to-SQL/query systems preferred.
- Experience with prompt engineering and prompt strategy development.
- Experience with GenBI or AI-driven analytics preferred.
- Experience with semantic models supporting natural-language querying.
- Experience with statistical analysis, predictive analytics, or optimization preferred.
- AWS AI services experience, including Bedrock and SageMaker, preferred.
Front-End / AI Application Skills – Preferred
- React development.
- Chainlit development.
- Streamlit development.
- UX/UI development for analytics dashboards.
- Highly interactive reporting applications.
- AI/ML or GenAI application development.
- Complex application state and data handling.
- Real-time updates.
- Loading, retry, error, and partial-response handling.
- Latency handling and asynchronous UX patterns.
- Explainable AI interfaces, including confidence indicators, sources, and disclaimers.
- AI chat interfaces with streaming responses and token handling.
- AI-assisted workflows such as:
- Autocomplete
- Summarization
- Recommendations
- Business insight generation
Certifications – Preferred
- AWS certifications.
- Snowflake certifications.
- BI certifications such as Power BI or Tableau.
- AI/ML or GenAI-related coursework, certifications, or credentials.
Education – Preferred
- Bachelor's degree in Computer Science, Computer Engineering, or a related field from an accredited or government-sanctioned college/university.
Key Deliverables
- ETL/ELT pipelines
- Snowflake data warehouse objects
- Data models
- Enterprise semantic layers
- APIs
- Data integrations
- GenBI solutions
- AI/ML and GenAI solutions
- Monitoring and logging dashboards
- Infrastructure automation
- CI/CD deployment solutions
- Technical documentation
- Knowledge-transfer materials
Project Success Metrics
- Pipeline reliability
- Data quality
- Query performance
- Platform uptime
- Deployment efficiency
- Successful delivery of investment-related GenBI/AI use cases
- Effective knowledge transfer to CalSTRS staff
Work Location Requirement
The selected candidate must be able to work in a hybrid environment, reporting onsite 2–3 business days per week at:
CalSTRS Headquarters
West Sacramento, California
Contract Information
- Initial Contract: Up to 12 months
- Extension Option: Additional 36 months
- Hours: Approximately 160 hours per month per person
- Payment: Hourly rate negotiated at contract execution
- Staffing: CalSTRS may select up to six technical resources
- Back-End Developers: Up to three positions
- Background Check: Required, including criminal, civil, and credit checks
Important Candidate Requirements
- Candidate must be able to work 2–3 days onsite per week in West Sacramento, CA.
- Candidate should have strong Data Engineering, Snowflake, SQL, ETL/ELT, and BI experience.
- Experience with AWS, GenAI, OpenAI APIs, semantic modeling, and GenBI is highly desirable.
- Experience with Airflow, dbt, Docker, Kubernetes, Terraform, and Ansible is required.
- Candidates with React, Chainlit, and Streamlit experience will be strongly preferred.
- Candidate must meet the required experience and education/certification requirements applicable to the position.
- Selected personnel must successfully complete the required CalSTRS background investigation.
- Candidate must be prepared to support knowledge transfer and documentation activities.
V Group Inc. is a NJ-based IT Services and Products Company with its business strategically categorized in various Business Units including Public Sector, Enterprise Solutions, Professional Services, Ecommerce, Projects, and Products. Within Public Sector business unit, we cater IT Professional Services to Federal, State and Local. We have multiple awards/ contracts with 30+ states, including but not limited to NY, CA, FL, GA, MD, MI, NC, OH, OR, CO, CT, TN, PA, TX, VA, NM, VT, and WA.
If you are considering applying for a position with V Group, or in partnering with us on a position, please feel free to contact me for any questions you may have regarding our services and the advantages we can offer you as a consultant.
Please share my contact information with others working in Information Technology.
Website: https://www.vgroupinc.com/publicsector
LinkedIn: https://www.linkedin.com/company/v-group/
Facebook: https://www.facebook.com/VGroupIT
Twitter: https://www.twitter.com/vgroupinc
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search