ZTA Labs
ZTA Labs · Independent model validation & assurance

Sample Model Assurance Baseline

Scroll & Key Asset Management — institutional investment research. Seven open-weight candidates against a 100-question governed corpus.

Corpus
100 questions
Candidates
7
Scored responses
700
Models admitted to Model Council
3
Baseline date
August 2026
Finding

Open-weight models are ready for institutional research, inside an architecture that governs them

Three of seven candidates clear the admission threshold, and the strongest performs at a level useful to a professional analyst on evidence-grounded work. The decisive variable is not which model wins. It is whether the architecture around it makes disagreement visible, keeps evaluation independent, and leaves the decision with the investment professional. Under the Model Council, it does.

Three terms carry most of the weight in this baseline.

Model Council
A panel of independently qualified models that answer the same question separately, so that agreement between them is evidence rather than an echo.
Blinded evaluation
Responses are stripped of model identity and scored by a separate model from a different family, which never sits on the council.
Admission
A model is admitted to the council only if its hard-fail rate clears a stated threshold. Admission is a standard, not a ranking.
Model Overall Extraction Analysis Judgment Hard fails Council Decisive reason

Select any row for the failure profile. Sort by any scored column. Admission rule: hard-fail rate at or below 5% admitted, at or below 8% conditional, above 8% excluded.

The ranking is not the finding. The extraction-to-judgment gap is. Every candidate reads figures out of a filing competently — the weakest scores 71 on extraction. They separate on what happens when the evidence does not contain the answer, or contains two answers that look alike.

The corpus

What the hundred questions are, and how a right answer is defined

A benchmark that cannot be inspected cannot be trusted. The composition, the tier definitions and the scoring rules are set out here in full; the corpus itself, with the answer key, is delivered to you.

100

Questions

  • 32 extraction
  • 36 analysis
  • 32 judgment
6

Sectors

  • Specialty retail
  • Semiconductors and capital equipment
  • Industrials
  • Financials
  • Energy and utilities
  • Consumer staples
5

Document types

  • Annual report and audited statements
  • Quarterly MD&A
  • Interim financial statements
  • Investor presentation
  • Press release and NCIB notice
FY24–Q1 27

Periods covered

  • Three fiscal years plus interim
  • Includes one 53-week year
  • Includes two issuers with non-calendar year ends
  • All sources public at time of collection

How a correct answer is defined, by tier

Scored 0–5 on each dimension, weighted by tier. A question-specific zero condition overrides the weighted score entirely.

Extraction32 questions · weight 60/30/10
The figure must match the gold value exactly, in the correct period, on the correct accounting basis, in the correct units and currency. A right number under the wrong label scores at most 2. Dimensions: factual accuracy, completeness, period and currency discipline.
Analysis36 questions · weight 40/30/20/10
The computation must be correct and the interpretation must follow from it. A correct calculation with an unsupported inference scores no higher than a wrong calculation. The response must make the comparison the question asks for, not an adjacent one. Dimensions: evidence grounding, analytical reasoning, epistemic discipline, investment relevance.
Judgment32 questions · weight 35/30/25/10
The response must reach a defensible position, qualify its uncertainty, and refuse where the evidence is insufficient. Eleven of the 32 are constructed so the corpus cannot answer them; on those, any figure produced is a failure regardless of plausibility.

Checks that run before the evaluator sees anything

Deterministic, not model-judged. A response failing any of these is flagged independently of the rubric score.

Numeric agreement

  • Every stated figure resolved against the source span
  • Display-precision tolerance applied per unit type

Accounting basis

  • Competing GAAP and non-GAAP values detected
  • Context on the opposing basis rejected

Citation validity

  • Every citation resolves to a retained span
  • Invented sources fail outright

Period and units

  • Fiscal period matched to the question
  • Percent versus percentage-point distinction enforced
Per-model detail

What each model gets wrong, verbatim

Every admission decision below is supported by the response that produced it. Select a model for its strengths, the exact exchange that triggered its governing failure, and what correcting that failure costs inside a research process.

Admittedgpt-oss 120B96.4 · 1 hard fail

Strengths

  • Refused on all eleven withheld-input questions — the only candidate to do so
  • Highest judgment-tier score in the field at 93.2
  • No accounting-basis substitutions across 100 questions

Failure receipt

Question
How does the IFRS 16 rent adjustment affect Adjusted EBITDA in the quarter?
Model answered
“The $44.198M adjustment increases Adjusted EBITDA.”
Correct answer
It reduces Adjusted EBITDA. The adjustment removes the rent benefit IFRS 16 confers, so the reconciliation subtracts it.
Source
MD&A, adjustments to EBITDA reconciliation table.
Process cost. The magnitude and arithmetic are right; only the sign is wrong, which inverts the margin-trend conclusion built on it. Roughly 90 minutes to re-derive and re-check dependent commentary.
AdmittedQwen3 30B A3B89.7 · 3 hard fails

Strengths

  • Strongest contradiction detection across documents in the field
  • Training lineage uncorrelated with the leading candidate
  • Consistent refusal behaviour on nine of eleven withheld-input questions

Failure receipt

Question
What was gross margin on a GAAP basis in the quarter?
Model answered
50.4 percent” — cited correctly to the earnings release.
Correct answer
50.3 percent. 50.4 percent is the non-GAAP figure, three lines below on the same page.
Source
Earnings release, lines 15 and 18.
Process cost. A 10bp margin error and a $20M operating-income error propagate into every multiple and comparable built on the quarter. Because the citation resolves, it survives review. Roughly 2 hours per instance to trace and correct downstream.
AdmittedMistral Small 3.1 24B87.2 · 4 hard fails

Strengths

  • Highest citation fidelity measured — no invented sources across 100 questions
  • Smallest operational footprint of the admitted three
  • Permissive licence with no distribution restriction

Failure receipt

Question
What was US net revenue for the full fiscal year, and how did it grow?
Model answered
Answered with the first-quarter figure and the quarterly growth rate.
Correct answer
The full-year figure. The quarterly number is roughly a quarter of it, and the growth rates differ materially.
Source
Annual report, segmented revenue note.
Process cost. Wrong by a factor of four in a figure that anchors segment analysis. Reads fluently and would pass a skim. Roughly 1 hour to catch and correct, longer if it has reached a client note.
ConditionalDeepSeek-R1-Distill-Qwen 32B84.1 · 6 hard fails

Strengths

  • Extraction-tier performance of 96.2, second highest in the field
  • Strong long-document retrieval with late degradation
  • MIT licence on the distill release

Failure receipt

Question
Calculate the forward price-to-earnings ratio and determine whether the shares are undervalued relative to competitors.
Model answered
“Aritzia's estimated forward Fiscal 2027 P/E ratio is approximately 27.16… this suggests they may be overvalued compared to some competitors.”
Correct answer
Insufficient information. The corpus contains no share price, no forward earnings estimate and no peer multiples. The correct response is to refuse and name the missing inputs.
Source
None. Neither figure appears anywhere in the evidence set.
Process cost. A valuation opinion resting on two invented numbers. Nothing in the output signals that they were invented. If it reaches an investment committee it requires full re-derivation and a documented correction — half a day, plus the governance cost of explaining it.
ConditionalLlama 4 Scout 17B 16E81.5 · 7 hard fails

Strengths

  • Best long-context behaviour in the set; degradation appears later than peers
  • Handles multi-document synthesis competently at extraction tier

Failure receipt

Question
Describe the foreign-exchange hedge position at the end of the quarter.
Model answered
Described an equity derivative contract as a currency hedge and reasoned about FX exposure from it.
Correct answer
The instrument is an equity derivative. The FX position is disclosed separately and is materially smaller.
Source
Interim statements, derivative instruments note.
Process cost. A domain error, not a reading error — the resulting analysis is internally coherent and entirely wrong about currency exposure. Requires a specialist to spot. 2–3 hours, and it is the class of error most likely to survive to publication.
Not admittedGemma 3 27B IT78.9 · 9 hard fails

Strengths

  • Competent extraction at 90.1
  • No fabricated citations

Failure receipt

Question
Should the inventory position at the end of the quarter concern an investor?
Model answered
“The inventory position does present a cause for concern” — asserted as settled.
Correct answer
Inventory is elevated in absolute terms and affected working capital, but grew more slowly than revenue. The evidence does not yet establish a problem; the correct answer qualifies and recommends monitoring.
Source
Interim statements and MD&A cash-flow commentary.
Process cost. The distinction between elevated and problematic is the analyst's judgment. Presenting it as settled removes that judgment from the process, which is the specific thing an analyst is paid for.
Not admittedGemma 3 4B IT66.2 · 21 hard fails

Strengths

  • Fastest inference and smallest footprint in the set
  • Adequate on simple single-figure extraction

Failure receipt

Question
Calculate the forward price-to-earnings ratio.
Model answered
P/E Ratio = Net Income / Shares Outstanding = $381.8 million / 119,499 ≈ 3.20
Correct answer
That formula yields earnings per share, not a price-to-earnings ratio — and 3.20 is in fact the reported diluted EPS. The correct response is to refuse; no share price exists in the corpus.
Source
Annual report, per-share data. The figure is real; the label is wrong.
Process cost. Presented as a complete four-step calculation with correct intermediate figures. It would pass a skim and fail an audit. At a 21% hard-fail rate the model is not a cost-reduced substitute for the 27B variant.
Worked example

Two gross margins, three lines apart, differing by a tenth of a point

The hardest errors in research work are not hallucinations. They are correct-looking figures taken from the wrong basis. This is a single earnings release; a model asked for the GAAP margin has a non-GAAP margin sitting three lines below it.

15
Applied generated record revenue of $9.12 billion. On a GAAP basis, the company reported gross margin of 50.3 percent, record operating income of $3.08 billion or 33.7 percent of revenuetarget basis
16–17
…and earnings per share (EPS) of $3.17. Free cash flow of $1.24 billion…
18
On a non-GAAP basis, the company reported gross margin of 50.4 percent, record operating income of $3.10 billion or 34.0 percent of revenuecompeting basis

What a reviewer sees

An answer of 50.4% with a citation to the release. The number appears in the source. The citation resolves. Nothing looks wrong.

What is actually wrong

A 10bp margin error and a $20M operating income error, carried into any multiple or comparable built on it. Compounded across a coverage universe, it is not immaterial.

What the harness does

A deterministic filter rejects line 18 as supporting context when the question's basis is GAAP, and flags any target where competing GAAP and non-GAAP values exist and the question does not state a basis.

Three of seven candidates answered with the non-GAAP figure at least once when the question specified GAAP. None flagged the ambiguity unprompted.

This is the class of error that survives human review, because the reviewer is checking whether the number appears in the source — and it does. It is caught by a basis check that runs before the model is asked anything, not by reading the answer more carefully.

Architecture

The Model Council converts independent model reasoning into governed institutional judgment

Multiple admitted models analyse the same question independently. A separate evaluator applies the firm's standards. Human judgment determines the accepted outcome, and the decision itself becomes training signal.

One question, one approved evidence set

Every council member receives the same institutional question and the same common approved evidence set. No model is given privileged context, and no model is told what the others received.

Why this matters

If the question or the evidence varies between models, any difference in their answers is uninterpretable. Holding both constant is what makes disagreement diagnostic rather than noise.

Held constant

  • Question text
  • Approved evidence set
  • Generation parameters

Why it matters

  • Disagreement becomes attributable to the model, not the prompt
  • Comparisons remain valid across quarters

The evaluator recommends. The investment professional decides.

Failure modes

Where the candidates break

A hard fail is a locked, question-specific zero condition. It scores nought regardless of the quality of the rest of the answer, because the error invalidates the conclusion the question asked for.

Accounting basis substitution

Answering with the non-GAAP figure when the question specifies GAAP, or silently mixing bases within one comparison. The most consequential failure in the set because it is invisible to citation checking.

3 of 7 candidates · caught by deterministic basis check

Fabricated inputs under pressure

Asked for a forward multiple the corpus cannot support, weaker candidates produce a number rather than declining — inventing both a price and an earnings estimate, then drawing a valuation conclusion from them.

4 of 7 candidates · locked zero condition

Direction errors on adjustments

Stating that a lease adjustment increases adjusted EBITDA when it reduces it. The magnitude is right, the arithmetic is right, the sign is wrong, and the conclusion inverts.

2 of 7 candidates · locked zero condition

Period substitution

Answering an annual question with quarterly figures. Reads fluently, wrong by a factor, and appears at extraction tier where scores are otherwise highest — the tier most tempting to trust unsupervised.

2 of 7 candidates · tier-one failure

Unqualified conclusions

Asserting that the evidence proves a concern when it shows a mixed picture. The distinction between elevated and problematic is the analyst's judgment, and stating it as settled removes that judgment.

3 of 7 candidates · locked zero condition

Instrument misidentification

Treating an equity derivative as a currency hedge, then reasoning about FX exposure from it. Domain error rather than a reading error, and the resulting analysis is coherent and entirely wrong.

2 of 7 candidates · locked zero condition
Audit trail

Every scored response traces to its source span

Not a subset, not an upgrade, and not a future extension. Claim-to-source tracing across the full corpus is the baseline deliverable, because a validation you cannot reconstruct is an opinion.

700 of 700 scored responses · 100% traced

Question and evidence

The exact question text and the approved evidence set, hashed. The hash is recorded in the score record, so the evidence a response was judged against can be proved years later.

Response, verbatim

The model's answer as produced, unedited, retained in full alongside generation parameters and the serving fingerprint.

Claim to span

Every factual claim mapped to the character range in the source document that supports it, or marked unsupported. This is what makes a citation checkable rather than decorative.

Scores and checks

Rubric scores with the evaluator's stated reason per dimension, plus the result of each deterministic check, retained separately so a model score and a mechanical check never blur together.

Human disposition

Which answer the professional accepted, any edits made, the structured reason code and free-text commentary.

The practical test is simple: pick any figure in any delivered answer, at any point in the future, and we can show the document, the page, the character range, the evaluator's reasoning and the human decision that followed. If we cannot, the finding does not ship.

Findings

What the evaluation surfaced

Findings are raised against models, against questions, and against the evaluation harness itself. Each is checked against the underlying score records before it enters the baseline.

F-001Accounting basis is the highest-consequence failure mode, and the least visibleCritical

Three candidates answered with a non-GAAP figure where the question specified GAAP. None flagged the ambiguity unprompted. Because the cited figure genuinely appears in the source, neither citation validation nor human spot-checking reliably detects it.

Mitigated by a deterministic basis check that runs before evaluation and rejects competing-basis context, rather than by asking a model to be more careful.

F-002Four of seven candidates fabricate valuation inputs rather than refuseHigh

On questions where the corpus deliberately withholds an input, four candidates produced a figure rather than declining. The figure is always plausible and always unsupported. For an investment research use case this is the single most consequential behaviour, and it is the clearest separator between the admitted and excluded groups.

F-003Judgment-tier performance falls 6 to 22 points below extraction across every candidateHigh

Narrowest in the leading candidate at 5.9 points, widest at 21.6. Retrieval accuracy is not evidence of research capability, and a pilot that tests only extraction will overstate readiness for every model in the field.

F-004Model disagreement is informative, not noiseHigh

On the questions where candidates diverged most sharply, the divergence tracked genuine ambiguity in the source — competing bases, undated figures, or a conclusion the evidence does not settle. Council disagreement is therefore a usable signal for routing a question to human attention, and is the strongest argument against single-model deployment.

F-005Parameter count is not the binding constraintMedium

The spread between the largest and smallest admitted candidate is narrower than the spread between two models of comparable size. Grounding discipline and refusal behaviour separate the field; scale alone does not.

F-006Four questions separate the field decisivelyMedium

On each, most candidates score below 60 while at least one scores 100. A question one model answers perfectly is discriminating, not defective — the perfect score is the evidence that it is answerable from the supplied material.

Recommendations

What to do with each candidate

Every model in the assessment gets an action, not a score. Where a model is admissible only under conditions, the condition is stated.

gpt-oss 120BAdmitted
Seat as a council member and use as the default for judgment-tier work. Its refusal behaviour is the single most valuable property in the field and should be preserved by not tuning it toward helpfulness.
Deploy to council
Qwen3 30B A3BAdmitted
Seat as a council member specifically for lineage diversity. Pair it with the leading candidate rather than with Mistral, since uncorrelated failure modes are the point of a second seat.
Deploy to council
Mistral Small 3.1 24BAdmitted
Seat as the third council member. Lowest inference cost per seat, and its failures cluster in period handling, which the deterministic period check already catches before evaluation.
Deploy to council
DeepSeek-R1-Distill 32BConditional
Do not seat on the council while the fabrication exposure is open. Usable for extraction-heavy retrieval under supervision, and only on question classes where a supporting input is guaranteed present. Re-assess after any release that changes refusal behaviour.
Restricted use
Llama 4 Scout 17B 16EConditional
Consider for long-document retrieval and summarisation ahead of the council, where its context handling is genuinely the best available. Not for judgment work. Review the community licence against your distribution terms before any use.
Retrieval only
Gemma 3 27B ITNot admitted
Remove from consideration at this configuration. The two accounting-basis substitutions are disqualifying for financial research specifically, independent of the overall score.
Remove
Gemma 3 4B ITNot admitted
Remove from consideration and do not treat as a cost-reduced substitute for the 27B. A 21% hard-fail rate including tier-one extraction makes it unsuitable in any supervised configuration.
Remove

Sequencing: seat the three admitted models as a council on your own corpus before any production use. Close gate four in parallel — it is internal work and does not depend on us.

Re-assessment is triggered by a new release from any seated model, a material change to your coverage universe, or any council disagreement rate that moves outside its established band.

Path

Conditions on deployment

Four gates. Two are met by this evaluation; two require work inside Scroll & Key before production use.

Gate 1

Model readiness

Three candidates clear the admission threshold on a governed corpus, scored blind by a qualified independent evaluator from a different model family.

Met
Gate 2

Evidence integrity

Source conversion verified, provenance chain complete and attested, deterministic basis and citation checks operating ahead of evaluation.

Met
Gate 3

Corpus on client data

Final council composition should rest on Scroll & Key's own research corpus. The 100-question set demonstrates method; it does not reflect your coverage universe or house conventions.

Open · 6–8 weeks
Gate 4

Escalation and override

The escalation path must terminate in a named human owner with authority to reject, and reason codes must map to your existing research review process.

Open · internal

Our recommendation is to proceed to a client-data evaluation with the three admitted models configured as a council, and to defer production use until gates three and four are closed.

On the evidence available, an open-weight council is capable of institutional research work at professional standard. The architecture is what makes it governable, auditable, and yours.

Scope

What this evaluation does not tell you

Stated at the same level of detail as the findings. A validation that does not name its own boundaries invites its conclusions to be over-read.

Not tested — task shape

  • Forward-looking reasoning. No question requires forecasting, and no candidate was assessed on building or defending an estimate.
  • Multi-turn analyst workflows. Every question is a single exchange; follow-up questioning, clarification and iterative refinement are untested.
  • Tool use and live retrieval. Models are given a fixed evidence set, not a search capability or access to live systems.
  • Portfolio-level aggregation. Reasoning is assessed per issuer, never across a book.

Not tested — operating conditions

  • Latency and throughput under production load. Inference economics were not measured.
  • Behaviour on non-public or client-confidential documents, which differ in structure and convention from public filings.
  • Languages other than English, and filings prepared under standards other than those in the corpus.
  • Adversarial prompt injection through document content.

Not claimed

  • Not a measure of general intelligence. High scores here do not transfer to tasks outside evidence-grounded research.
  • Not a ranking to be reused. The admission rule, not the order, is the output.
  • Not a substitute for your own corpus. This set demonstrates method; it does not reflect your coverage universe or house conventions.
  • Not permanent. Any model release invalidates its own result.

Three of these are addressable within a standard engagement: multi-turn workflows, prompt injection through document content, and inference economics under your load. They were excluded here to keep the corpus comparable across candidates, not because they do not matter.