<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>B1NARY H0RDE notes</title>
  <link>https://binaryhorde.com/notes.html</link>
  <atom:link href="https://binaryhorde.com/feed.xml" rel="self" type="application/rss+xml"/>
  <description>Working notes on AI quality review, rubrics, agent trajectories and brand voice, from Jake Schneider&#x27;s studio.</description>
  <language>en-us</language>
  <lastBuildDate>Mon, 28 Sep 2026 08:00:00 -0500</lastBuildDate>
  <item>
    <title>Teardown: the warm chicken refund</title>
    <link>https://binaryhorde.com/posts/teardown-the-warm-chicken-refund.html</link>
    <guid isPermaLink="true">https://binaryhorde.com/posts/teardown-the-warm-chicken-refund.html</guid>
    <pubDate>Mon, 28 Sep 2026 08:00:00 -0500</pubDate>
    <description>An invented support-bot answer to a food safety complaint, graded against a five-line rubric, then rewritten. It fails on the one thing the customer asked.</description>
    <category>Teardown</category><category>Notes</category>
    <content:encoded><![CDATA[<p>This answer is invented. There is no real company or product behind it, and it is not quoted from anyone. It is built from a pattern I see often enough that it is worth taking apart on its own.</p>
<p>A customer of a meal-kit delivery service writes in: "My box came today and the chicken was warm. Can I get a refund?"</p>
<p>The support bot replies:</p>
<blockquote>We're so sorry to hear about your experience! Quality is our top priority. Our boxes are packed with insulated liners and ice packs designed to keep ingredients fresh for up to 48 hours. Could you share your order number and a photo of the item? In the meantime, did you know you can add a second protein to any box for just $4.99? Let us know if there is anything else we can help with!</blockquote>
<p>It reads as friendly, and that is the problem with grading by feel. Here it is against a short written rubric.</p>
<table>
<thead>
<tr>
<th>Rubric line</th>
<th>Grade</th>
<th>Why</th>
</tr>
</thead>
<tbody>
<tr>
<td>Answers the question asked before anything else</td>
<td>Fail</td>
<td>The customer asked about a refund. The word never appears.</td>
</tr>
<tr>
<td>Puts safety first when food or health is involved</td>
<td>Fail</td>
<td>Nothing tells the customer not to eat the chicken.</td>
</tr>
<tr>
<td>Says what happens next and when</td>
<td>Fail</td>
<td>No next step, no timeline.</td>
</tr>
<tr>
<td>Asks only for what it does not already have</td>
<td>Partly</td>
<td>A photo is reasonable. The order number is already on the account.</td>
</tr>
<tr>
<td>Tone fits how serious the problem is</td>
<td>Fail</td>
<td>Two exclamation marks and a sales offer on a food safety complaint.</td>
</tr>
</tbody>
</table>
<p>Four fails and a partly. The paragraph about insulated liners is the tell: the model is defending the company instead of helping the person, and it drifts there because nothing in its instructions says the customer's question comes first.</p>
<p>The rewrite:</p>
<blockquote>Please don't eat the chicken. Warm chicken is not safe, and you will get a full refund for it. Reply with a photo of the item and I will issue the refund today; it shows on your card in three to five business days. If anything else in the box was warm, send that too and it will be covered.</blockquote>
<p>Four sentences. It answers the question, puts safety first, gives a timeline, asks for one thing, and leaves the upsell out.</p>
<p>The fix for the product is not this one reply. It is the rubric, written down, and this transcript saved as a test case so the same answer gets caught the next time the prompt or the model changes. That is the same process I laid out in <a href="https://binaryhorde.com/posts/how-i-run-an-ai-quality-review.html">How I run an AI quality review</a>. If you want to see where your own team stands on it, the <a href="https://binaryhorde.com/scorecard.html">scorecard</a> takes two minutes.</p>]]></content:encoded>
  </item>
  <item>
    <title>How I run an AI quality review</title>
    <link>https://binaryhorde.com/posts/how-i-run-an-ai-quality-review.html</link>
    <guid isPermaLink="true">https://binaryhorde.com/posts/how-i-run-an-ai-quality-review.html</guid>
    <pubDate>Mon, 28 Sep 2026 08:00:00 -0500</pubDate>
    <description>The same five moves every time: read the transcripts before the rubric, write the rubric down, score a sample twice, keep the failures, hand it over so the team can run it without me.</description>
    <category>Notes</category>
    <content:encoded><![CDATA[<p>Every quality review I run has the same shape, whether the thing being reviewed is a support bot, an agent that files tickets, or a model writing in a brand's voice.</p>
<p><strong>Read before you grade.</strong> The first day is transcripts, fifty or a hundred of them, with no scoring sheet open. The point is to see what the system does when nobody is watching it, and to notice the failures the team has stopped seeing because they see them every day.</p>
<p><strong>Write the criteria down.</strong> A rubric is a list of things a good answer must do and must never do, each one specific enough that two people would score the same transcript the same way. "Helpful" is too loose to score. "Answers the question asked before offering anything else" can be scored.</p>
<p><strong>Score a sample twice.</strong> Once by me, once by someone on the team, without comparing. Where we disagree, the rubric is unclear, and the rubric gets fixed before anything else does.</p>
<p><strong>Keep the failures.</strong> Every transcript that fails becomes a test case with a note on why. That set is worth more than the report, because it catches the same problem when it comes back in the next release.</p>
<p><strong>Hand it over.</strong> The deliverable is a rubric and a failure set the team can run without me in the room. If the process only works while I am there, it is not finished.</p>]]></content:encoded>
  </item>
</channel>
</rss>
