Gate B8: transcript entity-name accuracy has never been measured, yet names are being published #48

Open
opened 2026-08-23 17:33:10 -04:00 by logan · 1 comment
Owner

Board minutes #42 (2026-08-23) - Ratify BUSINESS_MODEL.md. - Gate B condition B8. Blocks charging any customer.

BUSINESS_MODEL.md §8 item 6, and the document own note on it: "this is the one people skip. You are about to sell machine-generated assertions about real incidents."

No measurement of transcript accuracy exists in source; the CISO (#39) confirmed it cannot be checked from a static read. We are currently publishing machine-extracted person and place names with an unknown error rate.

To close: sample 200 calls and hand-score (a) word error rate and (b) entity-name accuracy, separately. The second number is the one that matters.

Pre-committed now, so the result cannot be argued with after the fact (minutes #42 Gate B, B8): if entity-name accuracy is below 80%, names are suppressed by default. Committing to the threshold before seeing the number is the entire point.

Costs an afternoon and one spreadsheet, per the document. Requires human judgement - not agent-automatable.

Owner: the owner (human). Date: not set until the owner schedules it - see minutes #42 "Open".
Related: BUSINESS_MODEL.md §8 item 6, §9 Q3; the Gate B3 redaction issue.

Board minutes #42 (2026-08-23) - Ratify BUSINESS_MODEL.md. - **Gate B condition B8. Blocks charging any customer.** `BUSINESS_MODEL.md` §8 item 6, and the document own note on it: *"this is the one people skip. You are about to sell machine-generated assertions about real incidents."* No measurement of transcript accuracy exists in source; the CISO (#39) confirmed it cannot be checked from a static read. We are currently publishing machine-extracted person and place names with an unknown error rate. **To close:** sample **200 calls** and hand-score (a) word error rate and (b) **entity-name accuracy**, separately. The second number is the one that matters. **Pre-committed now, so the result cannot be argued with after the fact (minutes #42 Gate B, B8): if entity-name accuracy is below 80%, names are suppressed by default.** Committing to the threshold before seeing the number is the entire point. Costs an afternoon and one spreadsheet, per the document. Requires human judgement - not agent-automatable. Owner: **the owner (human).** Date: not set until the owner schedules it - see minutes #42 "Open". Related: `BUSINESS_MODEL.md` §8 item 6, §9 Q3; the Gate B3 redaction issue.
Author
Owner

Remit changed (not removed) by owner ruling 3, board minutes #42, "Owner rulings, 2026-08-23" comment.

The owner declined E&O spend pre-revenue ("assume it will happen but get to product mvp, feedback, then go for company things"). Gate B7's pre-agreed consequence is that person names are suppressed by default on every surface, paid included, until a policy is bound - see #43.

Therefore the pre-commitment in this issue - "if entity-name accuracy is below 80%, names are suppressed by default" - is now moot as a decision rule: names are suppressed either way. This issue is not dropped from Gate B. Its purpose changes to three things that still have to be true before we charge anyone:

  1. We are selling machine-generated assertions and have never measured their error rate. The "machine-generated, unverified" label (Gate A A2, #46) is a claim about accuracy we cannot currently characterise at all.
  2. The number is what an insurer will ask for at underwriting, and what would justify reinstating names once a policy is bound. Without it, the E&O revisit at company formation has no evidence to work from.
  3. Word error rate feeds the correlation quality argument independently of names.

Method unchanged: sample 200 calls, hand-score word error rate and entity-name accuracy separately. Still requires human judgement; still owner-scheduled.

**Remit changed (not removed) by owner ruling 3, board minutes #42, "Owner rulings, 2026-08-23" comment.** The owner declined E&O spend pre-revenue (*"assume it will happen but get to product mvp, feedback, then go for company things"*). Gate B7's pre-agreed consequence is that **person names are suppressed by default on every surface, paid included, until a policy is bound** - see #43. Therefore the pre-commitment in this issue - *"if entity-name accuracy is below 80%, names are suppressed by default"* - **is now moot as a decision rule**: names are suppressed either way. This issue is **not** dropped from Gate B. Its purpose changes to three things that still have to be true before we charge anyone: 1. We are selling machine-generated assertions and have never measured their error rate. The **"machine-generated, unverified"** label (Gate A A2, #46) is a claim about accuracy we cannot currently characterise at all. 2. The number is what an insurer will ask for at underwriting, and what would justify **reinstating** names once a policy is bound. Without it, the E&O revisit at company formation has no evidence to work from. 3. Word error rate feeds the correlation quality argument independently of names. Method unchanged: sample **200 calls**, hand-score word error rate and entity-name accuracy **separately**. Still requires human judgement; still owner-scheduled.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: logan/server-26#48