Insights
What we've learned building this.
Notes on the parts of AI engineering that decide whether a feature survives contact with real users. Opinionated, and specific enough to argue with.
Insights
Notes on the parts of AI engineering that decide whether a feature survives contact with real users. Opinionated, and specific enough to argue with.
Most teams building LLM features have no way to tell whether a change made things better or worse. Here is how to build the evaluation set that fixes that, and why it has to come first.
Fine-tuning teaches a model how to behave. Retrieval tells it what is true. Confusing the two is the most expensive mistake in AI project scoping.
New writing lands here as we hit problems worth documenting. No newsletter, no signup form — just a feed you can put in whatever reader you already use.
Subscribe via RSSIf you're making these decisions on a real product and want a second opinion, we're happy to talk it through — whether or not it turns into a project.
You can also read how we workfor how this thinking shows up in an actual engagement.
Start a conversation