Notes on agent evaluation, reproducible pipelines, and the plumbing that makes benchmarks trustworthy.
Evals are the unit tests of AI. Build and run them on Benchwright. Ship AI on scores, not vibes.
One email a month. New posts, eval results, and the occasional postmortem. No tracking pixels, no "just checking in."
We'll send the monthly dispatch to your inbox. One email a month — nothing else.