Thirty to fifty labelled cases over the vertical corpus, following docs/eval-golden-set.md exactly so they run in the same harness with the same scoring. Cover what the wedge depends on: extraction from document prose and from tables, verification against document spans, contradiction detection across documents including numeric and unit conflicts, supersession across document revisions, anchoring of subject entity plus document class plus revision, and retrieval and answering with correct citation of document locations.\n\nLabelling is the hard part and the value of the whole exercise. Write the labelling rules down BEFORE labelling: what counts as a fact worth extracting from a specification, what counts as a genuine contradiction versus a legitimate difference in scope or condition or precision, when two documents about different models must not be paired, and how to treat statements that are conditional or hedged in the source. Ambiguous cases are labelled conservatively and flagged, never silently decided. The rules live in the corpus documentation so a future labeller is consistent with this one.\n\nInclude negative cases deliberately, since precision is what a findings report lives on: same-boilerplate documents for different subjects that must not merge or contradict, numeric differences that are legitimate, superseding revisions that must supersede rather than contradict, and questions whose answers are genuinely not in the corpus. Every case stays traceable: source document, exact span, and the reasoning for its label.
Thirty to fifty labelled cases over the vertical corpus, following docs/eval-golden-set.md exactly so they run in the same harness with the same scoring. Cover what the wedge depends on: extraction from document prose and from tables, verification against document spans, contradiction detection across documents including numeric and unit conflicts, supersession across document revisions, anchoring of subject entity plus document class plus revision, and retrieval and answering with correct citation of document locations.\n\nLabelling is the hard part and the value of the whole exercise. Write the labelling rules down BEFORE labelling: what counts as a fact worth extracting from a specification, what counts as a genuine contradiction versus a legitimate difference in scope or condition or precision, when two documents about different models must not be paired, and how to treat statements that are conditional or hedged in the source. Ambiguous cases are labelled conservatively and flagged, never silently decided. The rules live in the corpus documentation so a future labeller is consistent with this one.\n\nInclude negative cases deliberately, since precision is what a findings report lives on: same-boilerplate documents for different subjects that must not merge or contradict, numeric differences that are legitimate, superseding revisions that must supersede rather than contradict, and questions whose answers are genuinely not in the corpus. Every case stays traceable: source document, exact span, and the reasoning for its label.