Baseline model qualification
One corpus, a defined model pool, firm-specific scoring, and an initial recommendation set.
Model Assurance
ZTA independently benchmarks open-weight models against the institution’s actual use cases and standards, then maintains the evidence, governance, and evaluation framework required to determine which models remain fit for use as models and institutional requirements change.
Why it matters
The initial baseline is only the beginning. Models change, the institution’s use cases evolve, and governance obligations grow more demanding as AI moves deeper into production work.
Prove model performance before production. Keep proving it after deployment.
One corpus, a defined model pool, firm-specific scoring, and an initial recommendation set.
Developments in the model pool, the benchmark, or operating landscape that matter to the client.
Evidence-backed recommendations on which models remain approved, restricted, or removed.
Review whether the methodology, controls, and audit record remain fit for purpose.
Out-of-cycle triage when a material release, new corpus, or new use case changes the decision surface.
What you get
Not a ranking. Each candidate is assessed on extraction, analysis, and judgment separately, because a model that reads figures flawlessly may still be unfit for the work that matters.
The corpus, answer key, scoring rules, and evaluation record remain in the client environment, ready to be re-run against future model candidates.
Every score traces back to evidence, hashed and attested. A regulator or risk committee should be able to inspect a record, not rely on a recollection.
The wrong basis, correctly cited. A model is asked for GAAP gross margin and returns the non-GAAP figure found a few lines lower in the same release.
The number is in the source, the citation resolves, and nothing reads as obviously incorrect. The error propagates quietly into every multiple built on it.
An unvalidated model does not announce these errors. It produces work that reads well and is wrong in ways that compound quietly across a coverage universe.
Sample baseline
A public example of the baseline work product: model-level findings, failure analysis, auditability, governance requirements, and deployment recommendations.
Independence
The model that judges the work must not be the model that produced it, and should not share its training lineage. Correlated failure modes mean a model can mark its own blind spots as correct.
We also treat the evaluator itself as something to be verified rather than assumed. Assurance that has never found a fault in its own process has not looked hard enough.