Methodology review
We read the model as a technical system: assumptions, target definition, feature construction, validation design, calibration, and the link between statistical evidence and the business decision.
Service · Independent review
StatGazer independently reproduces model results for investment teams before capital or governance depends on them. The handoff is a findings memo, issue register, reproducible evidence, and remediation path.
High-level context is enough. NDA before confidential data; scope, fee, and timing are agreed in writing.
Illustrative · synthetic data · not a real client engagement · not a track record
Model & Backtest Audit — Findings Memo
A deliberately flawed synthetic strategy is rebuilt on point-in-time inputs so the decision-maker can see what survives.
| Metric | Reported | Reproduced |
|---|---|---|
| Sharpe (annualized) | 1.8 | 0.7 |
| Annual return | 14% | 6% |
| Max drawdown | −9% | −21% |
| Hit rate | 58% | 52% |
ConfirmedHigh
A fundamental feature was joined on report date rather than its as-available date.
Need an independent first read before committing to a full validation? Model Preflight — $4,500 fixed, 5–10 business days.
Founder-led validation
Founder · Direct technical reviewer from scoping through handoff
When it matters
A model can look strong in research and still hide different classes of risk: unavailable features, weak conceptual assumptions, brittle calibration, undocumented monitoring, or evidence that cannot be reproduced. A useful validation separates those issues instead of collapsing them into a generic pass/fail view.
We read the model as a technical system: assumptions, target definition, feature construction, validation design, calibration, and the link between statistical evidence and the business decision.
We test whether inputs were available at the claimed decision time, whether joins are point-in-time, and whether universe construction, survivorship, or revision history contaminates the result.
We trace the path from source data to output. The goal is not a polished slide; it is evidence your team can regenerate, inspect, and defend under technical review.
How we work
The review follows a disciplined model-risk pattern: conceptual soundness, outcomes evidence, monitoring, limitations, and documentation. SR 26-2 — the Federal Reserve's revised model risk guidance, which superseded SR 11-7 in April 2026 — is useful vocabulary for that discipline, used here as a practical parallel rather than a claim of regulatory assurance.
Current supervisory vocabulary
The Federal Reserve's revised model risk guidance, SR 26-2 (April 2026, superseding SR 11-7), puts it plainly: "Model validation evaluates whether models perform as expected and includes an assessment of a model's reliability and its limitations." It also spells out what makes a challenge credible: "Effective challenge is performed by individuals with the appropriate expertise to conduct a critical and objective challenge, sufficient independence to maintain objectivity, as well as the organizational standing and influence to effect any change." We are not a bank supervisor's validator and do not claim regulatory assurance — but that definition, that expertise, and that independence are exactly what a Preflight and a full validation are built to deliver.
Topics below are commonly requested in institutional due-diligence questionnaires (AIMA-style DDQs and their equivalents). StatGazer documents evidence; it does not certify compliance with any questionnaire.
| Findings memo section | Questionnaire topic it typically answers |
|---|---|
| Reproduction attempt and evidence | Research process — backtest methodology and reproducibility of reported results |
| Data and timing checks | Data — sources, point-in-time integrity, look-ahead and revision controls |
| Validation design review | Model risk — validation approach, out-of-sample and walk-forward testing |
| Issue register with severity and decision impact | Model risk — known limitations, open issues, remediation status |
| Reproducible evidence package and handoff notes | Operations — documentation, change management, version control of models |
| Independent review itself | Governance — frequency and independence of model review |
High-level context is enough. NDA before confidential data; scope, fee, and timing are agreed in writing.
We validate models we did not build. If StatGazer designed or materially modified a model, we say so and recommend an independent reviewer for its validation. When an allocator engages us, we take no fee from the manager under review. The fee is fixed and does not depend on what we find; findings are reported as found, and there is no "pass" product. Remediation after findings is separate, optional work — and any later re-validation of a model we helped remediate is disclosed as such in the findings note.
Review evidence
Commercial fit
The strongest use case is not “we need a consultant.” It is “someone will ask whether this model is real, and our internal team needs a clean technical record before that conversation.”
Scope the review
Share the model type, the decision it informs, and your main concern. We will review the fit and agree the scope, fee, and timing before work begins.
Working inputs
Before we request code or data, we define the model use, the decision it supports, the review audience, and the level of evidence that would change the decision. That keeps the engagement focused and prevents a broad technical fishing expedition.
If the model is still early, the right engagement may be a narrower methodology review. If the model is already in production or near an allocation decision, the scope usually needs stronger evidence: reproducibility, out-of-sample behavior, failure modes, and a remediation path your team can execute. The point is to make the review decision-useful, not merely comprehensive.
A short model overview, the business decision, current validation evidence, known concerns, and the deadline or review event the work has to support. Confidential artifacts can wait until NDA and scope are agreed.
The review package ties each finding to supporting artifacts: code paths, data checks, validation outputs, screenshots, or documentation gaps. That keeps the memo inspectable after the readout.
The final readout is written for both technical owners and decision makers. Your team should leave knowing what was checked, what changed the conclusion, what to fix first, and what should be monitored after deployment.
Professional boundary. Technical model review, validation, research, and engineering consulting — not financial-statement audit, regulatory assurance, investment, legal, or tax advice.
Next step
Send a high-level description. Do not include confidential data, code, credentials, portfolio holdings, or client names until an NDA and written scope are in place.