Document Extraction Runs
Produced 31 July 2026
Where this came from
Six commercial lease documents filed with the SEC by REGENXBIO Inc., including a complete lease and a four-amendment chain
A batch where four of the six documents genuinely do not contain most of the field list. That is the case an extraction run is most likely to fail dishonestly, by inferring a value from a neighbouring document rather than reporting silence.
Twelve of twelve fields correct, every cite verified, and every blank audited.
The first run scored 66.7 percent against a 95 percent threshold. Every gain after that came from a defect an audit caught rather than from tuning: a character class that could not cross a hyphen, so money inside a parenthetical was unreachable; a non-greedy capture whose shortest legal match was the empty string; a label that recurs inside an estoppel-certificate template, which produced a populated, confidently-cited, completely wrong value; a table of contents mistaken for section headings; and two rounds of blanks sitting over stated dollar amounts. The scorer itself was replaced after it passed a wrong value on a loose substring match.