B1NARY H0RDE home status: taking work
ai_quality / scorecard 10 questions, about 2 minutes

AI quality scorecard

Ten plain questions about how your team checks what your AI product says to customers. Answer yes, partly, or no. The score is worked out in your browser and nothing is sent anywhere unless you choose to send it at the end.

the questions

Answer for how things work today

Not how they will work after the next planning cycle. Partly means it happens sometimes, or for some of the product.

  1. Before anyone decided what a good answer looks like, did someone read a real batch of transcripts, fifty or more, with no scoring sheet open?

    Reading first is how you find the failures the team has stopped noticing.

  2. Is there a written rubric: a list of what a good answer must do and must never do?

    If the standard lives in people's heads, it changes with whoever is reviewing.

  3. Is each item on that rubric specific enough that two people would score the same answer the same way?

    "Helpful" cannot be scored. "Answers the question asked before offering anything else" can.

  4. Have two people scored the same sample separately, without comparing notes, and then checked where they disagreed?

    Disagreement between two reviewers is the fastest way to find a vague rubric line.

  5. When reviewers disagree, does the rubric get rewritten before anything else changes?

    Fixing the model against an unclear rubric just moves the problem.

  6. Is every failed answer saved as a test case, with a note on why it failed?

    That saved set of failures is worth more than any single report.

  7. Before each release, do you rerun that set of saved failures to check nothing old came back?

    Old problems return quietly when a prompt or model changes.

  8. Is one named person responsible for the tone of what the product says and whether customers can trust it?

    When everyone owns tone, nobody does, and it drifts.

  9. Are tone, trust, brand voice, and clarity scored with the same weight as factual correctness?

    In a product that talks to customers, how it says something is part of whether it worked.

  10. Could your team run the whole review next month without the person who set it up?

    If it only works while one person is in the room, it is not finished.

0 of 10 answered