My parser finds bugs in its own reading, without an answer key

8 August 2026

I have a parser that reads government documents and reports when two sources contradict each other about the same thing. It runs on Japanese disaster bulletins — one report says a town’s water is out, a later one says restored, and someone has to notice.

The parser has bugs. Every new document format produces a new one. What I could not work out for a long time was how to find them without me reading the output, because “is this reading correct?” needs somebody who knows what the document means.

Then I noticed I had been asking the wrong question. There is a different one that does not need the world:

not:  is this reading correct?
but:  do two readings of the SAME CONTENT agree?

If they disagree, one of them is wrong. That is a proof, not a heuristic, and it costs no human. This is metamorphic testing — the trick is finding a transform where you can argue the direction, not just the disagreement.

The transform

Japanese does not put spaces between words. So a space between two kanji in running prose was put there by the PDF extractor, not by the author. That gives you something stronger than “these two readings differ”:

LAYOUT CANNOT ADD INFORMATION.

If closing up an extractor’s space makes a claim disappear, the claim was manufactured by the whitespace. Not “suspicious” — spurious. No arrangement of spaces is evidence that a town has water.

Concretely, from a real ministry PDF:

「全 12 戸が断水しています」 →  parser reads: 全 ("all") is out of water
「全12戸が断水しています」   →  parser reads: 全12戸 ("all 12 households")

The first is a fragment. Nobody had to read either one to know that one of them is wrong.

What it actually found

Run across five corpora of real published documents:

13  proven defects on two ministry PDF series
 0  on statutes, municipal HTML, operator press releases

Not typos. Things like a claim about water restoration filed under 自治体 (“municipality”, the generic word) instead of the actual municipality’s name, and a service disruption filed under 路線 (“route”) instead of the line name. Anyone asking about their own town or their own train line got nothing back.

The repair is mechanical, because the answer key is internal

Once you can prove a defect, you can propose a fix and measure it. The gate:

the planted test suite still passes
no confirmed finding across five corpora is lost
coverage does not fall
the count of proven defects strictly falls

Two candidates, and both outcomes happened — which is why both code paths exist:

counter_split  ACCEPTED
  A numeral and its counter are one word (12戸, 15炉).
  Proven defects 13 → 12, coverage 73.39% → 73.39%,
  the same 9 confirmed findings, the same 18,460 sentences placed.

layout_space   REJECTED
  Close up ANY single space between two CJK characters.
  Removes every proven defect — and costs 79 sentences their
  subject, of which only 8 were the spurious claims.

The rejected one is the more useful record. Without it, the same losing candidate gets proposed on every single run, forever. So the rejection goes into a ledger.

A second oracle: the output versus the parser’s own rules

The parser has guards that mean “if this pattern follows the term, the term asserts nothing”: 〜のため (“for the purpose of”), 〜による (“caused by”), 〜と認める (“deemed to be”).

So a placed claim whose tail one of those guards matches is an internal contradiction. Both the output and the rules live in the same process — no world knowledge enters. That found 7 more:

「災害復旧のため派遣された職員」
   →  filed "restored" on 災害派遣手当 ("disaster dispatch allowance")

A dispatch allowance is not a restored water main.

And this is the class I had hit four separate times by hand: a guard applied on the prose path and skipped on the table path. Enumeration, deeming, until, and now のため. Every time, a human found it. So I fixed the class instead of the instance — suppressions are now consulted at the one line every claim passes through, which means the hole cannot reopen as a path-skip.

Where it stops

This does not make the parser self-improving in any general sense. It repairs what the parser’s own reader broke. It cannot tell you what a word it has never seen MEANS — no transformation of a document reveals that — so new vocabulary arrives as a queue with an approve button, and nothing in that path can write to the config without a person pressing it.

The honest summary: metamorphic relations gave me an answer key for the class of bug where the input was misread, and nothing at all for the class where the parser is simply ignorant. That turned out to be a bigger fraction than I expected, and a smaller one than I wanted.

Numbers, because a post like this is worthless without them

corpusfindingstrue
government disaster reports, 5 corpora, 4 read blind1414
naive keyword baseline, same documents386
technical prose, 93 mixed EN/JA documents50

That last row is on the front page of the README. The engine works where documents make state claims about NAMED things — a municipality, a route, a contract. On prose it manufactures contradictions out of abstract nouns that recur in unrelated contexts, and a wiki is the wrong input.


Code: github.com/Ag3497120/cleanroom · pip install verantyx-vera

No LLM anywhere in the answer path, no GPU, runs offline. That last part is not a flex — the people this is for have shelter registers and hospital lists on their laptops, and the correct place for those is nowhere.