ML Engineer, Image Systems
Indexed description
Models·Remote
- Pacific time overlap·Full-time contractor·Competitive compensation
Make image generation and editing work reliably inside HawtAds. You will choose models, catch their failure modes, and own the tradeoffs between fidelity, speed, and cost.
Apply for this role →Questions first? →
Make frontier models dependable for ad work.
HawtAds helps marketing teams generate, edit, and ship ad creative. A model can look impressive in a benchmark and still fail the work: change the product, break the headline, ignore a reference, or turn a precise edit into a full rerender.
Your job is to make those systems dependable inside HawtAds. You will test them on real ad work, choose where each belongs, fix what can be fixed, and know when to move on. You will work closely with product engineers, because a great model behind a clumsy workflow is still a bad product.
What you own.
- Choose the right generation or editing approach for each workflow. Defend the default with real quality evidence, latency, reliability, and cost.
- Build the evals that matter for ads: readable type, faithful products and people, good reference use, useful composition, precise edits, and policy compliance.
- Improve weak spots through model-specific prompting, reference conditioning, routing, and lightweight adaptation when the evidence supports it.
- Make model calls boring in production with sensible routing, fallbacks, retries, rollout controls, policy checks, and cost limits.
- Work with product and editor engineers on the parts customers feel: controls, progress, failure recovery, and the way one edit leads to the next.
- Keep up with new releases and silent behavior changes without chasing every launch into production.
This is roughly how the responsibility should grow. We expect useful work early, then wider ownership as you learn the product and customers.
- Week 1 to 2 Make one model-backed workflow measurably better. Watch customer sessions, learn the current quality bar, and find a real failure our evals miss.
- Month 1 Take over the evals for one workflow and recommend a model, routing, prompting, or adaptation change. Show us the quality, latency, and cost case.
- Month 3 The default strategy for at least one workflow is yours, with a measurable improvement in quality, reliability, speed, or unit economics.
- Month 6 You own the image-systems roadmap, from evals and adaptation to production behavior and the way new capabilities enter the product.
- 3+ years shipping ML systems in production, including being responsible for what happens after launch.
- You have worked hands-on with current multimodal image systems, not only classic text-to-image diffusion. Native editing, multi-reference conditioning, masks, text rendering, and consistency across iterations should all be familiar ground.
- You can turn “this looks better” into a useful dataset, review process, launch threshold, and monitoring plan.
- You have made hard quality, latency, reliability, and cost tradeoffs on real traffic and can explain the call clearly.
- Strong Python and PyTorch skills, plus comfort working at a typed product API boundary.
- You write clearly enough that product, design, and engineering can understand a decision months later.
- Parameter-efficient adaptation or fine-tuning for products, brands, people, or other identity-sensitive image domains.
- Multimodal evaluation work, including ranking data, calibrated model judges, and blinded human review.
- Inference optimization for open-weight models, including batching, quantization, compilation, and GPU scheduling.
- Trust and safety, age-gated workflows, or platform-policy compliance for generative imagery.
- Experience in advertising, design tools, or another product where model output quality is the user experience.
The frontier will move before this posting comes down. We care less about loyalty to one model family than your ability to evaluate the next one rigorously.
Generation and editing Native multimodal systems for creation, precise edits, and iterative workflows
Reference control Multi-image conditioning, product and identity consistency, and edit locality
Adaptation Parameter-efficient methods when prompting and routing are not enough
Evaluation Task-specific golden sets, blinded pairwise review, and calibrated multimodal judges
Delivery Hosted APIs when they win; self-hosted open weights when they earn the extra work
Product quality Type people can read, products and people that stay recognizable, and edits that change the right thing
How hiring works.
We do not run a gauntlet. Each step should give both of us useful signal without unnecessary delay.
- Apply Send a CV and a short note about one production image-system decision you made. Tell us what the evidence said, what you traded off, and what happened. No cover letter.
- Intro 45 minutes with the founder. We will talk about what you would own, why now, and whether the role makes sense on both sides.
- Deep dive A take-home exercise, designed for about 6 hours, based on a real image-quality problem.
- Final conversations Two remote sessions: a review of the exercise with the team and a system-design conversation about evaluating and delivering image capabilities in a product.
- References + offer Two references of your choosing. An offer follows within a week of the final conversations.
Start with one production image-system decision you can explain with evidence. That is what we read first.
Apply for this role →← Back to all roles
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search