A bid package becomes a risk register, and three stages of it do not run
Left to right: PDFs go in, a ranked register comes out. The label on each arrow is the data structure that crosses it. The fill on each box says whether that stage runs today.
- runs today
- built and tested, wired to nothing
- not built
- input or data file
Wide diagram. Drag it sideways on a narrow screen.
- Ingest runs on real PDFs. It pulls text out of a package, splits it by CSI MasterFormat section numbering, and keeps the document, page and rectangle for every piece. No model is involved.
- Scope marker detection runs too, and its gap is recorded. It finds phrases such as "NIC", "by others" and "furnished by" using patterns, not a model. It misses owner-furnished scope written in ordinary English, and that miss is held open as a backlog item rather than papered over.
- Claim extraction is finished and reaches nothing. This is the single sanctioned model call for reading documents. It is written and tested, but no command invokes it, so no package the tool can actually produce today contains claims.
- Two-vendor agreement is in the same state. Complete, tested, called by nothing outside its own tests. Diagram 3 is what it does when it is wired.
- Plan extraction does not exist. Nothing in the system reads a drawing sheet. Spatial models are written by hand today, which is why lane B is dotted up to that point.
- Everything meets here, and this half is deterministic. The rule engine cannot reach a model at all. Not the vendor libraries, not an HTTP client, not even the packages that hold them. Diagram 2 is how that is enforced.
- The register is the product. The bill of materials is supporting evidence under it, not the headline, and the rule trace records which rules were suppressed as well as which fired.
- The viewer reads the JSON, never the report. It renders values the engine already computed. It performs no geometry and no scoring of its own, so it cannot disguise a missing engine capability.
Two consequences of the marked gaps are worth stating. Because claim extraction is unwired, half of the seam analysis in diagram 4 has nothing to read on any package this tool can produce today. And because there is no benchmark, none of this says the findings are correct.
A model may say what a document says. Only code may say what that means.
One wall, drawn down the middle. The wall is a schema. Nothing downstream of it ever reads model text.
- Everything on the left is reading. The model finds a sentence, says where it is and what it appears to say. It is used the way you would use a very fast reader with no authority.
- The wall is a schema, not a policy. What the model is permitted to return has no page, no rectangle and no section number on it. Those fields do not exist on the record it fills in, so it cannot supply them.
- Code opens the gate by finding the quote on the page. It searches the actual page for the exact text the model returned. Found means the rectangle is built from the search result. Not found means the reading is dropped and the quote is counted as a defect, which is what stops invented text from acquiring a citation.
- Everything on the right is a determination. A quantity, a part number, a severity, a compliance call, and a citation are all conclusions about the world, so a model may not make any of them. Rules written in versioned data files make them instead.
- The wall is checked by machine, because the failure mode is a reasonable person under deadline. Someone will want one model call inside a validator to resolve an ambiguity, and they will be right that it would work. A review convention depends on a reviewer noticing on that day. An import graph does not.
The reason for the split is commercial rather than aesthetic. You bring a flag to a pre-bid review and the engineer whose design you are questioning asks where it came from. A named rule, a quoted sentence and a page number turns that into a discussion about the document, which is the discussion worth having.
Two vendors read the same row, and there is deliberately no tiebreak
Two paths out of one comparison. One of them produces a value. The other produces no value at all, on purpose.
- One row, two readers, no contact between them. Neither model is shown the other's answer, because two readers who can see each other are one reader.
- A row only one vendor read is a disagreement, not a single-vendor value. The absent reader contributes an explicit "not read" at zero confidence. Skipping it would let a row one model never saw arrive looking like a well-read row.
- The comparison is code, and it is held to the engine's standard. No model, no network, no clock. A test parses the comparator and fails if it imports a vendor library, an HTTP client, or anything that could make two runs differ.
- Agreement takes the lower confidence, never a higher one. Two readers each eighty percent sure does not become ninety-six percent sure. A formula that lifts them would be a number this system invented about its own reliability. Agreement also does not rescue a bad reading: agreeing on an illegible cell is still agreeing on an illegible cell.
- Disagreement emits nothing, and the record has no winner field. Two independent readers diverging on a construction document is evidence about the document, not about the models. It says the passage is genuinely ambiguous, which is exactly what an estimator is paid to know before the bid closes.
- Why a tiebreak is the wrong feature. A flag can be ignored by a later step that forgot to check it; an absent value cannot. And a tiebreak shortens the flag list without reducing the ambiguity, so the estimator reads the shorter list as good news about the documents rather than as a change in the software.
A run refuses to start when only one reader is configured. Quietly falling back to a single vendor would report agreement that was never checked, and the register would come out wrong in the direction of false confidence.
One door opening, two divisions, and four ways the money goes missing
Door hardware is Division 08, written by the architect. Access control is Division 28, written by the security engineer. The electrified locking device sits on the line between them.
- The seam is where the failure is symmetrical. A sentence reads "by others, by Division 28". The security contractor reads it as somebody else's scope. The door hardware supplier reads the same sentence, agrees the work is Division 28, and neither party prices the rough-in.
- Both, and they disagree. Two sections carry the device and want different things from it. That argument happens at submittal, after the price is fixed. If they carry it and agree, it is a duplicate buy instead, which is a different rule and a different pre-bid question.
- Neither, and both sections were read. This is only reportable because the engine knows it saw both sides. Saying nobody specified it on a package that never included Division 08 would be a confident answer built on a missing document.
- The circular hand-off is its own state, not an error. Computed naively it looks identical to "both divisions carry it", which is the opposite finding and the opposite action. It carries a higher severity and an action that says to quote both sentences side by side in the request for information, because asking with only one attached invites the answer that it is obviously yours.
- Not determined is an admission about our reading, not a conclusion about the documents. It exists so that state has somewhere to go. A bidder shown "nobody specifies the rough-in" on a package whose Division 08 section was never supplied will find the assignment themselves and stop believing the register.
- The healthy case produces nothing, and that is measured. A Division 08 section handing door contacts to Division 28 is how a well-drafted specification is supposed to read. If the silence ratio ever inverts, the rules have started training their reader to skip the section.
The full seam analysis reads two kinds of evidence: located scope statements, which the ingest lane produces today, and structured claims, which it does not. Until claim extraction is wired, half of this runs on nothing. That is stage 3 in diagram 1.
What "every line carries provenance" means, on one line
Demonstration. The line below came out of an engine run against the repository's golden fixture: an invented package for a fictional district. The district, the manual, the page and the sentence are all fabricated for testing. It is here to show the shape of a finding and where each part comes from. It is not evidence that the engine is right, and no accuracy claim is made anywhere on this page.
- The rule id and its pack are on the line itself. Two runs a month apart can be compared because the rule that fired is named in both, and pack ids are permanent. They are never renumbered.
- The citation is looked up by code, not written by a model. A citation is a claim about the world, so the record is unconstructable without at least one locator. A line without provenance is treated as a defect in the software, not a formatting preference.
- The sentence is quoted verbatim. A tidied reading is what a reviewer ends up arguing with instead of arguing with the document.
- Rationale and authority travel together. Code interpretation varies by jurisdiction, so a rule cites what it leans on and stays overridable on a project. The rationale is written as if the reader is the engineer whose design you just contradicted, because sometimes they are.
- The action is phrased for the week before the bid closes. A pre-bid request for information is free. The same question asked after award is a negotiation.
- Confidence is carried, never smoothed. Below the threshold the item is flagged for a person. Where the documents do not determine something, the engine emits a flag rather than a guess.
- The exposure band is a placeholder and says so. Every band shipped today was chosen so the ranking key has something to sort on, and none is derived from cost history. Rather than print a placeholder in the same typeface as a real figure, every rendering marks it. No figures appear on this site for that reason.
- The rules that lost are recorded too. When two rules could apply to one sentence, one wins and the other is kept with its reason, so a reviewer who expected a different finding can read the rule that displaced it instead of guessing. Reporting both would put the same money in the register twice.
A run replays exactly from the same documents, configuration and seed, down to the individual characters of the report. That matters for an ordinary reason and an unusual one. If the register changes, something changed and you can find out what. And a run that cannot be replayed cannot be benchmarked, which is the thing this project does not have yet.