Scroll & Key Asset Management — institutional investment research. Seven open-weight candidates against a 100-question governed corpus.
Three of seven candidates clear the admission threshold, and the strongest performs at a level useful to a professional analyst on evidence-grounded work. The decisive variable is not which model wins. It is whether the architecture around it makes disagreement visible, keeps evaluation independent, and leaves the decision with the investment professional. Under the Model Council, it does.
Three terms carry most of the weight in this baseline.
| Model | Overall | Extraction | Analysis | Judgment | Hard fails | Council | Decisive reason |
|---|
Select any row for the failure profile. Sort by any scored column. Admission rule: hard-fail rate at or below 5% admitted, at or below 8% conditional, above 8% excluded.
The ranking is not the finding. The extraction-to-judgment gap is. Every candidate reads figures out of a filing competently — the weakest scores 71 on extraction. They separate on what happens when the evidence does not contain the answer, or contains two answers that look alike.
A benchmark that cannot be inspected cannot be trusted. The composition, the tier definitions and the scoring rules are set out here in full; the corpus itself, with the answer key, is delivered to you.
Scored 0–5 on each dimension, weighted by tier. A question-specific zero condition overrides the weighted score entirely.
Deterministic, not model-judged. A response failing any of these is flagged independently of the rubric score.
Every admission decision below is supported by the response that produced it. Select a model for its strengths, the exact exchange that triggered its governing failure, and what correcting that failure costs inside a research process.
$44.198M adjustment increases Adjusted EBITDA.”50.4 percent” — cited correctly to the earnings release.first-quarter figure and the quarterly growth rate.approximately 27.16… this suggests they may be overvalued compared to some competitors.”equity derivative contract as a currency hedge and reasoned about FX exposure from it.does present a cause for concern” — asserted as settled.P/E Ratio = Net Income / Shares Outstanding = $381.8 million / 119,499 ≈ 3.20”The hardest errors in research work are not hallucinations. They are correct-looking figures taken from the wrong basis. This is a single earnings release; a model asked for the GAAP margin has a non-GAAP margin sitting three lines below it.
An answer of 50.4% with a citation to the release. The number appears in the source. The citation resolves. Nothing looks wrong.
A 10bp margin error and a $20M operating income error, carried into any multiple or comparable built on it. Compounded across a coverage universe, it is not immaterial.
A deterministic filter rejects line 18 as supporting context when the question's basis is GAAP, and flags any target where competing GAAP and non-GAAP values exist and the question does not state a basis.
Three of seven candidates answered with the non-GAAP figure at least once when the question specified GAAP. None flagged the ambiguity unprompted.
This is the class of error that survives human review, because the reviewer is checking whether the number appears in the source — and it does. It is caught by a basis check that runs before the model is asked anything, not by reading the answer more carefully.
Multiple admitted models analyse the same question independently. A separate evaluator applies the firm's standards. Human judgment determines the accepted outcome, and the decision itself becomes training signal.
Every council member receives the same institutional question and the same common approved evidence set. No model is given privileged context, and no model is told what the others received.
If the question or the evidence varies between models, any difference in their answers is uninterpretable. Holding both constant is what makes disagreement diagnostic rather than noise.
Council size varies by workflow and risk. No model sees another model's response, so agreement is evidence rather than an artefact of one model anchoring the others.
Models shown each other's work anchor to the first answer they see. Isolation is what prevents a single confident error from becoming unanimous agreement.
Responses are stripped of model identity and scored blind against the firm's rubric. The evaluator is qualified independently and is never a council member — a model does not grade work it could have produced.
A model that grades work it could have produced marks its own blind spots correct. Separating the evaluator by family is the difference between review and self-assessment.
The analyst does not receive a single answer. They receive the shape of the disagreement, which is the part that carries information.
A single merged answer hides the disagreement that should have prompted a second look. Showing the divergence is what turns the council into a routing signal for analyst attention.
Four dispositions, each recorded. A structured reason code accompanies every decision, which is what makes the resulting signal usable rather than anecdotal.
Without a structured reason, a rejection teaches nothing. The reason code is what converts one professional's judgment into something the firm can aggregate.
Each exchange is retained and linked in full. Over a year of use the firm accumulates thousands of recorded judgments — which answer a professional accepted, which they rejected, and why. That corpus improves retrieval, sharpens routing, tightens the rubric, and eventually trains a model on your house view specifically. It appreciates with use, and it belongs to you.
This record is the one asset in the stack that no vendor can supply and no competitor can copy. It is a longitudinal account of what your professionals judged correct, on your coverage, in your house conventions.
The evaluator recommends. The investment professional decides.
A hard fail is a locked, question-specific zero condition. It scores nought regardless of the quality of the rest of the answer, because the error invalidates the conclusion the question asked for.
Answering with the non-GAAP figure when the question specifies GAAP, or silently mixing bases within one comparison. The most consequential failure in the set because it is invisible to citation checking.
Asked for a forward multiple the corpus cannot support, weaker candidates produce a number rather than declining — inventing both a price and an earnings estimate, then drawing a valuation conclusion from them.
Stating that a lease adjustment increases adjusted EBITDA when it reduces it. The magnitude is right, the arithmetic is right, the sign is wrong, and the conclusion inverts.
Answering an annual question with quarterly figures. Reads fluently, wrong by a factor, and appears at extraction tier where scores are otherwise highest — the tier most tempting to trust unsupervised.
Asserting that the evidence proves a concern when it shows a mixed picture. The distinction between elevated and problematic is the analyst's judgment, and stating it as settled removes that judgment.
Treating an equity derivative as a currency hedge, then reasoning about FX exposure from it. Domain error rather than a reading error, and the resulting analysis is coherent and entirely wrong.
Not a subset, not an upgrade, and not a future extension. Claim-to-source tracing across the full corpus is the baseline deliverable, because a validation you cannot reconstruct is an opinion.
The exact question text and the approved evidence set, hashed. The hash is recorded in the score record, so the evidence a response was judged against can be proved years later.
The model's answer as produced, unedited, retained in full alongside generation parameters and the serving fingerprint.
Every factual claim mapped to the character range in the source document that supports it, or marked unsupported. This is what makes a citation checkable rather than decorative.
Rubric scores with the evaluator's stated reason per dimension, plus the result of each deterministic check, retained separately so a model score and a mechanical check never blur together.
Which answer the professional accepted, any edits made, the structured reason code and free-text commentary.
The practical test is simple: pick any figure in any delivered answer, at any point in the future, and we can show the document, the page, the character range, the evaluator's reasoning and the human decision that followed. If we cannot, the finding does not ship.
Findings are raised against models, against questions, and against the evaluation harness itself. Each is checked against the underlying score records before it enters the baseline.
Three candidates answered with a non-GAAP figure where the question specified GAAP. None flagged the ambiguity unprompted. Because the cited figure genuinely appears in the source, neither citation validation nor human spot-checking reliably detects it.
Mitigated by a deterministic basis check that runs before evaluation and rejects competing-basis context, rather than by asking a model to be more careful.
On questions where the corpus deliberately withholds an input, four candidates produced a figure rather than declining. The figure is always plausible and always unsupported. For an investment research use case this is the single most consequential behaviour, and it is the clearest separator between the admitted and excluded groups.
Narrowest in the leading candidate at 5.9 points, widest at 21.6. Retrieval accuracy is not evidence of research capability, and a pilot that tests only extraction will overstate readiness for every model in the field.
On the questions where candidates diverged most sharply, the divergence tracked genuine ambiguity in the source — competing bases, undated figures, or a conclusion the evidence does not settle. Council disagreement is therefore a usable signal for routing a question to human attention, and is the strongest argument against single-model deployment.
The spread between the largest and smallest admitted candidate is narrower than the spread between two models of comparable size. Grounding discipline and refusal behaviour separate the field; scale alone does not.
On each, most candidates score below 60 while at least one scores 100. A question one model answers perfectly is discriminating, not defective — the perfect score is the evidence that it is answerable from the supplied material.
Every model in the assessment gets an action, not a score. Where a model is admissible only under conditions, the condition is stated.
Sequencing: seat the three admitted models as a council on your own corpus before any production use. Close gate four in parallel — it is internal work and does not depend on us.
Re-assessment is triggered by a new release from any seated model, a material change to your coverage universe, or any council disagreement rate that moves outside its established band.
Four gates. Two are met by this evaluation; two require work inside Scroll & Key before production use.
Three candidates clear the admission threshold on a governed corpus, scored blind by a qualified independent evaluator from a different model family.
Source conversion verified, provenance chain complete and attested, deterministic basis and citation checks operating ahead of evaluation.
Final council composition should rest on Scroll & Key's own research corpus. The 100-question set demonstrates method; it does not reflect your coverage universe or house conventions.
The escalation path must terminate in a named human owner with authority to reject, and reason codes must map to your existing research review process.
Our recommendation is to proceed to a client-data evaluation with the three admitted models configured as a council, and to defer production use until gates three and four are closed.
On the evidence available, an open-weight council is capable of institutional research work at professional standard. The architecture is what makes it governable, auditable, and yours.
Stated at the same level of detail as the findings. A validation that does not name its own boundaries invites its conclusions to be over-read.
Three of these are addressable within a standard engagement: multi-turn workflows, prompt injection through document content, and inference economics under your load. They were excluded here to keep the corpus comparable across candidates, not because they do not matter.