Machine Language Engineer
Indexed description
Description:
We are looking for competence in designing, configuring, deploying, and maintaining machine learning solutions for anomaly detection and early warning in a data center environment. The solution must run on-premises and integrate with existing operational data streams, including InfluxDB as the time-series database and Kafka as part of the data pipeline. The role requires practical experience with machine learning, time-series analytics, Linux operations, and production deployment in industrial or infrastructure environments.
Required Competence
- Experience with time-series data analysis and anomaly detection for infrastructure systems such as cooling, power, UPS, environmental sensors, network equipment, and IT load.
- Ability to evaluate and select suitable ML models for anomaly detection, forecasting, and early warning use cases.
- Knowledge of statistical anomaly detection, Isolation Forest, Autoencoders, LSTM models, Prophet, ARIMA, streaming anomaly detection, and hybrid rule/ML-based systems.
- Ability to explain trade-offs related to accuracy, explainability, training data requirements, operational complexity, false positives, and real-time performance.
Programming and Development Competence
- Experience with recognized programming languages commonly used for machine learning and data engineering such as Python, Go, Java, Scala, Rust, or C++.
- Understanding of frameworks and tooling for machine learning pipelines, data processing, and model serving.
- Experience building APIs and backend services for operational environments.
- Ability to work with containerized and distributed systems.
Technical Environment
- Ubuntu Linux configuration, hardening, and operations.
- InfluxDB integration for time-series storage and retrieval.
- Kafka integration for streaming pipelines and event-driven architectures.
- Docker or Kubernetes-based deployment.
- REST API development using frameworks such as FastAPI, Spring Boot, Go services, or similar technologies.
- Logging, monitoring, alerting, and production operations.
Data Pipeline and Integration
- Reading real-time or near-real-time data from Kafka.
- Querying historical data from InfluxDB for model training.
- Writing anomaly scores and prediction results back to InfluxDB.
- Exposing results through APIs, dashboards, or Grafana integrations.
- Supporting integration with ITSM or operational monitoring systems.
Deployment Responsibilities
- Installing required dependencies and runtime environments.
- Configuring services using systemd, Docker, or Kubernetes.
- Managing model artifacts and versioning.
- Scheduling retraining and inference jobs.
- Implementing health checks, logging, and operational metrics.
- Documenting deployment, rollback, and maintenance procedures.
Operational Requirements
- Designing robust and explainable anomaly detection systems.
- Reducing false positives while ensuring meaningful early warning capability.
- Defining training data requirements and retention periods.
- Handling missing data, sensor errors, and seasonal operational patterns.
- Distinguishing between normal operational variation and real anomalies.
Desired Deliverables
- Recommended ML model architecture.
- Data pipeline design using Kafka and InfluxDB.
- Ubuntu deployment guide.
- Configuration files and service definitions.
- Model training and inference implementation.
- Alerting and anomaly scoring logic.
- Operational and maintenance documentation.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search