Knowledge extraction / May 2026

Wrong twenty times, useful anyway

I built a contradiction detector for a product knowledge base. Every single flag it raised was a false positive. Then the false positives turned out to have more structure than the truths, and the tool ended up measuring something better than what I built it for.

The setup: take a SaaS product's public help center, 665 articles, and extract every factual claim into subject, predicate, object triples. 2,292 claims. Then flag every place where two articles give the same subject and predicate a conflicting object. The dream is automated documentation QA: "article A says the export includes companies, article B says projects, one of you is wrong."

An earlier normalization pass had already taught me some humility. Simply canonicalizing surface forms cut the flag count from 64 to 12, an 81% reduction from string hygiene alone. What survived that pass, plus a corpus expansion, left 41 object-level mismatches across 20 predicate groups. Those 20 are the subject of this piece, because I sat down and hand-audited every one.

Claims
2,292
Flag groups
20
Real errors
0
Artifact classes
4

01 Zero for twenty

Not one flag was a genuine documentation error. Not one was even an intentional design split. All twenty were artifacts of the extraction itself: places where the pipeline had manufactured a disagreement out of text that was perfectly consistent.

That is a bleak headline for a tool whose name is "contradiction detector." The interesting part is what the audit found instead of contradictions. The twenty false positives were not a smear of random noise. They sorted cleanly into four classes, each with its own telltale, and each pointing at a specific, fixable extractor behaviour.

A. Partial enumeration 7 of 20

Same subject and predicate, but each article enumerates the subset of objects relevant to its own scope. A per-integration article lists the two siblings it interacts with; an overview article lists all of them. Every claim is independently true. The objects are disjoint but consistent.

Telltale: no article's claim excludes another's. Fix: compare union sets, and only flag when one article's enumeration rules out what another asserts.

B. Paraphrase split 5 of 20

The same fact wearing two outfits. "Next step" versus "next campaign step," where the subject is the campaign. "Center" versus "centered." Semantically identical under the lightest synonymy, but past the reach of string normalization.

Telltale: equivalence a human sees instantly. Fix: lemma-level object matching, and folding predicate tails when the subject already disambiguates them.

C. Scope drift 4 of 20

Multiple articles each describe a different facet of one capability, freshly phrased every time. An assistant feature "detects response rhythm" in one article, "detects deadlines" in another, "detects commitments" in a third. These are bullet points of one list, not rival claims. Especially common in articles transcribed from videos, where nothing is ever phrased the same way twice.

Telltale: all objects could sit in a single bulleted list under the feature. Fix: cluster objects by similarity within a predicate group before flagging.

D. Wrong span 3 of 20

The extractor grabbed the wrong words. "Relies on Apple software to recognize your phone" became relies on: phone, which then "contradicted" the correct relies on: Apple software from another article. The disagreement is between the extractor and itself.

Telltale: one object is grammatically downstream of the real one. Fix: dependency-aware span selection, or re-prompting the extractor with tighter span guidance.

02 The biggest class is not a bug

Three of the four classes are straightforward extractor defects with engineering fixes. Class A is different, and it is the largest.

Partial enumeration flags fire precisely where the documentation answers one question piecemeal across many scoped articles, with no single page holding the complete answer. The flag is wrong about a contradiction existing. It is exactly right about something else: this is a fact that has been fragmented. If seven scoped articles each carry a shard of "what does search cover," that is where a canonical reference page should exist and does not, and it is where the shards will drift apart the next time one article gets updated and six do not.

A docs team does not especially need a tool that says "these two pages disagree," because they rarely do. They could genuinely use a tool that says "this fact lives in seven places and nobody owns the whole of it."

So the honest description of what I built is not a contradiction detector with 0% precision. It is a fragmentation auditor with a misleading name. Same pipeline, same flags, different claim about what a flag means, and the second claim survives an audit while the first one went zero for twenty.

The tool was measuring something real the whole time. It was just mislabeled by its own aspiration. Renaming what a signal means is sometimes the entire fix.

03 What the audit bought

The uncomfortable counterfactual: without the hand audit, the headline would have been "41 contradictions found across the knowledge base," and it would have been fiction. Twenty sections was a few hours of reading. Every count a tool reports is a claim about the world until a human has sampled it, and the sample here returned the worst possible verdict for the headline and the best possible input for the roadmap.

Because a false-positive taxonomy is a roadmap. "The detector is wrong 100% of the time" is unactionable despair. "Seven flags are union-set comparisons, five are lemma equivalence, four are missing clustering, three are span errors" is a sprint plan, ranked by count, with an acceptance test attached to each class: after the fix, that class's flags should vanish and no others should.

What I would not claim

This is one corpus, one extraction pipeline, twenty flags. The four classes are certainly not a universal taxonomy of extraction errors; they are what this corpus surfaced, and a legal or medical corpus would presumably grow different ones. The base-rate lesson travels, though. In a professionally maintained knowledge base, real contradictions are rare, so a detector's flags will be dominated by its own artifacts almost regardless of how good it is. Any contradiction-flavoured tool should expect its early output to be a mirror held up to its extractor, and plan the audit accordingly.