Computer Scientist 2
Indexed description
Computer Scientist (preferably Level 2) – Need as soon as possible.
Preferred Experience
- End-to-end neural speech-recognition systems using wav2vec 2.0, Whisper, and Deep Speech 2-style acoustic models.
- Sequence-to-sequence and speech models for multilingual and domain-specific applications.
- Experience with curating real and synthetic datasets for training, validation, and model comparison.
- Building datasets for repeatable benchmarking and regression testing.
- Designing benchmark suites measuring model accuracy, inference latency, package size, and release readiness.
- Designing hybrid recognition architectures that combine local processing with remote API calls.
- Applying language models and NLP techniques to correct likely recognition errors and improve contextual interpretation.
- Building scalable ETL and data-enrichment pipelines using e.g. Python, PySpark, Rust, Go, and FastAPI.
- Developing production ML systems using PyTorch, TensorFlow, Keras, MLflow, ZenML, and related tools.
- Implementing artifact validation, deployment controls, and AI governance for production model releases.
- Applying anomaly detection, statistical modeling, and error analysis to identify operationally significant model failures.
- Rapid AI prototyping and translating research models into reliable, production-grade software pipelines.
CASPER duties:
· Design, implement, and maintain a centralized repository for CASPER training data, validation data, transcriptions, command labels, model outputs, and evaluation results.
· Provide full-stack, DevOps-oriented development across the training and validation environment, including front-end data-management interfaces, back-end services and APIs, databases, processing infrastructure, automated builds, deployment, monitoring, and operational support.4
· Develop aviation-specific model harness for LLM API integration
· Develop and train custom language models for simulated pilot speech and behaviors
· Build scalable ETL, labeling, and data-enrichment pipelines that map recorded and synthetic speech to CASPER commands, callsigns, fixes, digits, and other recognition targets.
· Maintain dataset, label, model, and command-schema versioning so training and evaluation data remain aligned as CASPER capabilities and phraseology mappings change.
· Curate representative training and held-out validation datasets covering different speakers, phraseology, audio conditions, callsign types, commands, and operational contexts.
· Automate end-to-end model training, replay, benchmarking, validation, and nightly or per-build regression testing of the CASPER recognizer and parser.
· Measure and report word error rate, callsign accuracy, digit accuracy, command and intent accuracy, precision and recall, latency, regressions, and release readiness.
· Use exploratory, hypothesis-driven, and experiment-based methods to identify, prototype, and evaluate alternative approaches for improving voice-recognition performance.
· Design controlled experiments and ablation studies comparing acoustic models, language models, contextual sources, preprocessing methods, confidence thresholds, and recognition-pipeline configurations.
· Develop, fine-tune, and evaluate ATC-specific speech-recognition models, including wav2vec 2.0, Whisper, Deep Speech 2-style models, sequence-to-sequence models, and other promising architectures.
Remote Work: Remote work may be permitted when consistent with project needs. On-site presence will be required as dictated by FAA project needs.
Experience and Qualifications:
Masters or Phd in Computer Science or related field preferred
Preferred experience in AI/ML
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search