Staff Software Engineer, Command Alert
Indexed description
Learn more about us here!
This is a 100% remote position
About Your New Role
Command|Alert is CommandLink's signal-processing core: the engine that turns raw security, monitoring, and customer-defined telemetry into alerts customers actually trust. Alert fatigue and noise are the top complaint across every competitor in this space, and this role exists to make sure our alerts are the ones people don't tune out.
As a Staff Software Engineer on Command|Alert, you'll operate across the two to three teams that touch alerting, from rule evaluation and anomaly detection through delivery and downstream notification. You'll drive the org's most consequential decisions on how the correlation and alerting engine is architected, dive into whichever team or project needs your depth, and balance near-term reliability work with the long-term technical foundation the product is built on.
This is a role for someone who reasons fluently across data: taking in security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols, deriving real network and system topologies and the dependencies between them, and using that context to make LLM-driven reasoning over correctness, troubleshooting, and remediation actually work.
Key Responsibilities
- Set the architecture for how Command|Alert evaluates rule-based thresholds, ML anomaly scores, and correlation logic to produce high-fidelity alerts, across both global alerts we define and alerts customers define themselves.
- Own the reliability of the alerting pipeline end to end: from OpenSearch alert evaluation through Kafka delivery via OpenSearch callbacks to downstream notification, including idempotency guarantees and soak-tested behavior under sustained load.
- Drive the correlation strategy that turns diverse sources (security tooling, monitoring telemetry, syslog, OpenTelemetry, L2-L4 network protocols) into usable network and system topologies and their dependencies, and make that context usable for LLM reasoning over investigations and remediation.
- Lay the technical groundwork for generating alert definitions from the normalized data model using LLMs.
- Jump into any team or workstream across Command|Alert that needs architectural guidance, unblocking others and raising the bar on how the system is built.
- Balance strategic bets (new correlation and detection capability) against the long-term foundation the alerting engine needs to hold up at scale.
- Mentor engineers across the teams you touch, and represent Command|Alert's technical direction to stakeholders outside engineering.
- Takes on additional responsibilities and projects as needed to support the success of the team and organization.
- Deep, hands-on experience with OpenSearch alerting and ML-based anomaly detection, including running that logic in production against real traffic and load.
- Strong Kafka experience, particularly around producing and consuming alert events for downstream notification.
- Experience with Temporal as a workflow orchestration client.
- A strong track record with webhook reliability and idempotency, and prior experience building or operating alerting systems in production.
- A working command of the protocols and data these systems consume, security tooling output, monitoring telemetry, syslog, OpenTelemetry, NetFlow/sFlow, SNMP, ICMP, and firewall logs, and how to turn that data into real topologies and dependency maps.
- Recognized mastery of Go and/or Python, with the range to work across stream processing (Flink), Kubernetes/Helm/Docker, and a multi-cloud footprint (AWS, Azure, GCP).
- A demonstrated ability to make org-level architecture and technology calls, not just execute within one.
- Experience using LLMs to reason over structured system or network data for troubleshooting, remediation, or investigation workflows.
- Familiarity with Protocol Buffers for service contracts, Argo/Spacelift for CI/CD and infrastructure-as-code, and Memgraph or another graph store for topology data.
- Background with osquery, Steampipe, or similar endpoint/cloud inventory tooling.
- Experience operating in a multi-tenant, cloud-native environment with secrets management and TLS at scale.
- Room to grow at a high-growth company
- An environment that celebrates ideas and innovation
- Your work will have a tangible impact
- Flexible time off
- Fun events at cool locations
- Employee referral bonuses to encourage the addition of great new people to the team
AI tools are used only to assist in the evaluation process — they do not make final hiring decisions. Every application is reviewed by a member of our recruiting or hiring team before any decisions are made.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search