Service · Independent review

Independent model validation.

StatGazer independently reproduces model results for investment teams before capital or governance depends on them. The handoff is a findings memo, issue register, reproducible evidence, and remediation path.

High-level context is enough. NDA before confidential data; scope, fee, and timing are agreed in writing.

StatGazer / Public sample work product Ref. SAMPLE-001

Illustrative · synthetic data · not a real client engagement · not a track record

Model & Backtest Audit — Findings Memo

A material part of the headline does not survive point-in-time reconstruction.

A deliberately flawed synthetic strategy is rebuilt on point-in-time inputs so the decision-maker can see what survives.

Reported versus reproduced metrics from SAMPLE-001 — fabricated synthetic data, not a client result or track record
MetricReportedReproduced
Sharpe (annualized)1.80.7
Annual return14%6%
Max drawdown−9%−21%
Hit rate58%52%

ConfirmedHigh

Look-ahead in the signal pipeline

A fundamental feature was joined on report date rather than its as-available date.

Read the complete sample memo (opens in a new tab)

Engagement profile

Best for
Models already influencing allocation, sizing, risk, or oversight decisions.
Output
Findings memo, reproducible evidence, issue register, and remediation plan.
Delivery
Founder-led · typically two to six weeks once scope and access are agreed.
Fee
From $15,000, scoped in writing.

Need an independent first read before committing to a full validation? Model Preflight — $4,500 fixed, 5–10 business days.

Evgenii Azarov, founder of StatGazer

Founder-led validation

Evgenii Azarov

Founder · Direct technical reviewer from scoping through handoff

  • Econometrics · statistics · machine learning · financial engineering
  • PhD in Law · Russia, 2012 · finance training at NYU
  • 20 years of private teaching and consulting · New York · global
Inspect the full credibility ledger →

When it matters

Validation creates a taxonomy of findings.

A model can look strong in research and still hide different classes of risk: unavailable features, weak conceptual assumptions, brittle calibration, undocumented monitoring, or evidence that cannot be reproduced. A useful validation separates those issues instead of collapsing them into a generic pass/fail view.

Methodology review

We read the model as a technical system: assumptions, target definition, feature construction, validation design, calibration, and the link between statistical evidence and the business decision.

Leakage and timing

We test whether inputs were available at the claimed decision time, whether joins are point-in-time, and whether universe construction, survivorship, or revision history contaminates the result.

Reproducibility

We trace the path from source data to output. The goal is not a polished slide; it is evidence your team can regenerate, inspect, and defend under technical review.

How we work

Independent challenge, written evidence.

The review follows a disciplined model-risk pattern: conceptual soundness, outcomes evidence, monitoring, limitations, and documentation. SR 26-2 — the Federal Reserve's revised model risk guidance, which superseded SR 11-7 in April 2026 — is useful vocabulary for that discipline, used here as a practical parallel rather than a claim of regulatory assurance.

What we examine

  • Data lineage, timestamp discipline, joins, missingness, and revision history.
  • Feature construction, target definition, model choice, assumptions, and parameter stability.
  • Validation design: train/test split, walk-forward testing, regime sensitivity, calibration, and benchmark comparisons.
  • Operational controls: versioning, monitoring, documentation, and handoff readiness.

Finding taxonomy

  • Confirmed defects with evidence and decision impact.
  • Material assumptions that are defensible only under stated conditions.
  • Unresolved questions caused by missing evidence, access, or documentation.
  • Remediation items ranked by model-use risk, not cosmetic neatness.

Current supervisory vocabulary

SR 26-2 on validation and effective challenge.

The Federal Reserve's revised model risk guidance, SR 26-2 (April 2026, superseding SR 11-7), puts it plainly: "Model validation evaluates whether models perform as expected and includes an assessment of a model's reliability and its limitations." It also spells out what makes a challenge credible: "Effective challenge is performed by individuals with the appropriate expertise to conduct a critical and objective challenge, sufficient independence to maintain objectivity, as well as the organizational standing and influence to effect any change." We are not a bank supervisor's validator and do not claim regulatory assurance — but that definition, that expertise, and that independence are exactly what a Preflight and a full validation are built to deliver.

Source: Board of Governors of the Federal Reserve System, SR 26-2, “Revised Guidance on Model Risk Management”, April 17, 2026.

Where the findings land in an institutional questionnaire

Topics below are commonly requested in institutional due-diligence questionnaires (AIMA-style DDQs and their equivalents). StatGazer documents evidence; it does not certify compliance with any questionnaire.

Findings memo section Questionnaire topic it typically answers
Reproduction attempt and evidence Research process — backtest methodology and reproducibility of reported results
Data and timing checks Data — sources, point-in-time integrity, look-ahead and revision controls
Validation design review Model risk — validation approach, out-of-sample and walk-forward testing
Issue register with severity and decision impact Model risk — known limitations, open issues, remediation status
Reproducible evidence package and handoff notes Operations — documentation, change management, version control of models
Independent review itself Governance — frequency and independence of model review

Confidentiality

High-level context is enough. NDA before confidential data; scope, fee, and timing are agreed in writing.

Independence

We validate models we did not build. If StatGazer designed or materially modified a model, we say so and recommend an independent reviewer for its validation. When an allocator engages us, we take no fee from the manager under review. The fee is fixed and does not depend on what we find; findings are reported as found, and there is no "pass" product. Remediation after findings is separate, optional work — and any later re-validation of a model we helped remediate is disclosed as such in the findings note.

Review evidence

What validation catches

Commercial fit

Built for the buyer who has to defend the answer.

The strongest use case is not “we need a consultant.” It is “someone will ask whether this model is real, and our internal team needs a clean technical record before that conversation.”

Good fit

  • A strategy backtest is moving toward allocation or increased sizing.
  • A risk or forecasting model is being used in a recurring decision process.
  • A PM, CIO, risk committee, IC, or vendor-review process needs defensible evidence.
  • The internal team wants an outside reviewer without handing away model ownership.

Not a fit

  • You need a financial-statement audit opinion or regulatory attestation report.
  • You want someone to bless a model without access to evidence.
  • You need investment advice, legal advice, tax advice, or a performance guarantee.
  • You want confidential data reviewed before an NDA and scope are in place.

Scope the review

Tell us what the model needs to support.

Share the model type, the decision it informs, and your main concern. We will review the fit and agree the scope, fee, and timing before work begins.

Send a model brief

Working inputs

Validation starts with context, not a file dump.

Before we request code or data, we define the model use, the decision it supports, the review audience, and the level of evidence that would change the decision. That keeps the engagement focused and prevents a broad technical fishing expedition.

If the model is still early, the right engagement may be a narrower methodology review. If the model is already in production or near an allocation decision, the scope usually needs stronger evidence: reproducibility, out-of-sample behavior, failure modes, and a remediation path your team can execute. The point is to make the review decision-useful, not merely comprehensive.

What helps upfront

A short model overview, the business decision, current validation evidence, known concerns, and the deadline or review event the work has to support. Confidential artifacts can wait until NDA and scope are agreed.

Evidence package

The review package ties each finding to supporting artifacts: code paths, data checks, validation outputs, screenshots, or documentation gaps. That keeps the memo inspectable after the readout.

Handoff standard

The final readout is written for both technical owners and decision makers. Your team should leave knowing what was checked, what changed the conclusion, what to fix first, and what should be monitored after deployment.

Professional boundary. Technical model review, validation, research, and engineering consulting — not financial-statement audit, regulatory assurance, investment, legal, or tax advice.

Next step

Start with the model and the decision it supports.

Send a high-level description. Do not include confidential data, code, credentials, portfolio holdings, or client names until an NDA and written scope are in place.

Send a model brief