Back to search
Whym Linkedin · Posted 2d ago

Founding Engineer, Search & Relevance

Bend, Oregon, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About Whym


Whym is an AI-powered app that helps people make time for what matters most. We suggest activities aligned with your values and help you coordinate plans with close friends. We are a two-person founding team with experience from Google, LinkedIn, Lyft, and earlier-stage startups. Whym is a Public Benefit Corporation with pre-seed funding.


About the role


You will own Whym’s recommendation quality end to end. Every suggestion we put in front of a user passes through systems you will build: the catalog pipeline that assembles what there is to do, the ranking that orders what we show, and the evaluation that tells us whether it works. Most of those systems do not exist yet. The mandate is scrappy: stand them up from first principles, capture the low-hanging fruit, and let eval help you find the next win.


Quality is central to the product’s promise. Users respond to the concept. Whether they stay depends on whether each suggestion lands.


This is a build-heavy role. Expect to spend most of your time shipping production code: services, pipelines, indexes, and harnesses, with modeling and analysis in service of what you ship.


This is a tiny startup. Everyone does customer support, including the CEO and CTO. You will too. Some weeks you will debug a ranker. Other weeks you will review user feedback, build a labeling UI, or sit with a tester to understand why a suggestion missed. You should find that appealing, not beneath you.


What you will work on


  • Ranking and relevance. Design, build, and ship the systems that decide what Whym suggests to whom, and in what order. Own the full stack: retrieval, candidate generation, ranking, and freshness. Iterate based on eval signals and user feedback. Today’s ranker is an LLM self-score placeholder. You will replace it with something learned, one measured win at a time.
  • Data pipeline and indexing. Build the infrastructure that turns the sprawl of internet-accessible sources (event listings, venue pages, local calendars) into a coherent, current view of what there is to do in the real world. You will design the index, decide what we own versus fetch on demand, and expand source coverage as we open new markets. Scraping, integration, and data partnerships run through this function.
  • Evaluation. Build the system that tells us whether a quality change helped or hurt. Design rubrics. Run human eval. Operate LLM-as-judge at scale. Maintain offline test sets. Set up AB tests. Create the metrics dashboards the team lives by. Without eval, quality work is guesswork. You will make sure we are not guessing.
  • Quality research. Answer the hard questions with data. Where are we failing? Which suggestions convert? What does a good session look like? Then build agentic workflows that ask those questions continuously: agents that sweep for failure patterns, propose ranking changes, and check them against eval. Research sits next to ranking and eval because all three only work when they feed each other.


What we’re looking for


  • 8+ years in engineering, with meaningful time on search, feed, recommender, ads, or ranking systems at consumer scale.
  • Strong software engineering foundations. You have designed, built, and operated production systems end to end: services, schemas, pipelines, and the infrastructure under them. Not only models and notebooks.
  • You have shipped ranking or retrieval systems that real users depend on. We want to hear about specific decisions you made, what you measured, and how the system improved over time.
  • Solid grounding in ranking and experimentation fundamentals. You can design an experiment, reason about metrics, and tell a good signal from a noisy one. You do not need to be a researcher.
  • Hands-on with evaluation. You know how to build test sets, write rubrics, and run both human and LLM-based eval. You have opinions about what breaks when you try to scale eval, and how to fix it.
  • AI-native. You use modern AI tools in your daily work. You have run LLM-as-judge pipelines and understand their failure modes. You adopt new techniques before they are mainstream.
  • Fluent with data infrastructure. You can design a pipeline, debug a schema, and argue for or against owning an index. You do not need a platform team to get work done.
  • Strong product instinct. You know how to measure success and which levers to pull to get there. You connect ranking decisions to user outcomes and can say why a metric matters in plain language.
  • Clear writer and communicator. You document decisions. You present results without spin. You disagree without posturing. You know what AI slop looks like and you do not turn it in as final work.
  • Comfortable with ambiguity. You thrive at early stage, where the definition of quality is itself evolving.
  • High agency. You notice what needs doing and do it without waiting to be asked.


Strong pluses


  • Experience building infrastructure for recommendation-driven products. You have seen how it gets built at places that do it well, and you are ready to build it again from first principles.
  • Time in ads or search relevance, where the daily work is measuring success and finding the levers that move it.
  • Background building evaluation frameworks from scratch, not just operating an existing one.
  • Experience with local discovery, location-based products, or anything involving fresh real-world inventory.
  • Public write-ups or talks about systems you built.
  • Experience managing agentic or LLM-driven systems in production.
  • A point of view on how modern AI changes what ranking and evaluation look like.
  • Experience at an early-stage startup where you wore many hats.
  • Systems thinker. You build processes that scale. You document what you build.


Location and working model


Whym is a hybrid team based in Oakland, California and Bend, Oregon. We work together in person three days a week and reserve the other days for focused remote work. We prefer candidates who can work with us at our hubs, but will make room for exceptional candidates who want to be fully remote. Bend-based and remote teammates travel to the Bay Area regularly so the whole team can spend time together.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search