Manager technical due diligence for family offices and allocators.
Before you allocate to a systematic manager, an independent reviewer reproduces what the deck claims — from the manager's code where it can be shared, from the documents where it cannot — and reports what holds, what does not, and what is still unresolved. You engage us; you pay us, per decision; the manager under review pays nothing.
Founder-led · One named reviewer · Fixed fee per decision, no subscription · No fee from managers · NDA first · Reply within 24 hours
No sales pitch. No obligation. We say no when the date or the materials will not support a useful read.
Built for the desk that has to sign off without the code.
Large single-family offices run diligence like institutions. Most family offices, OCIOs, and smaller funds of funds do not have a quant on staff — yet they are asked to approve allocations to managers whose case rests on a backtest, a model, or a track record they cannot reproduce. This page is for that desk.
Who engages us
Single- and multi-family offices without an in-house quant team.
OCIOs and investment consultants adding systematic or model-driven sleeves.
Funds of funds, seeders, endowments, and foundations at manager-selection or re-underwriting stage.
When it is worth a read
A committee date is set and the case rests on a backtest or model output.
A due-diligence questionnaire answers "research process" in prose you cannot verify.
An existing allocation is being resized, or the manager has moved to a new model version.
Live results have drifted from the backtest and nobody can say why.
What you get back
A one-page summary — decision, principal findings, what remains unresolved — written to be forwarded as is.
Evidence: what was reproduced, what was not, and the checks that were run.
A list of precise questions to put to the manager, in the manager's own terms.
Managers we read: systematic equity and futures, trend and CTA programs, machine-learning alpha, statistical arbitrage, options overlays, QIS providers, digital-asset and DeFi quant strategies, and the systematic sleeves of multi-strategy platforms — the technical layer only.
Screening, pre-allocation, and monitoring need different depths. Each engagement is bought per decision — no subscription, no minimum number of managers — and is fixed in scope and fee before work begins. The fee does not depend on what we find.
Compare three engagements across stage, question, inputs, checks, deliverable, timeline, fee, and limitation
Engagement
Desk Read
Model Preflight
Independent Model Validation
Stage
Desk ReadStageScreening, or when the manager will not share code
Model PreflightStagePre-allocation, one model
Independent Model ValidationStageAllocation already carrying weight, or a committee that needs the full record
Question it answers
Question it answersDo the documents hold together, and what should we ask before going further?
Question it answersIs there a material problem worth a deeper look?
Question it answersCan this model be trusted for the decision it supports?
What the manager provides
What the manager providesDocuments only: deck, backtest report, tear sheet, returns, trade list if available, DDQ answers on research process
What the manager providesCode at a pinned state, one data snapshot, and a working environment with read-only access
What the manager providesFull code, data lineage, documentation, and environment
What we do
What we doSix checks: recompute the headline statistics from the returns supplied; test statistical significance and multiple-testing exposure; check cost, capacity, and turnover plausibility; compare backtest with live where live exists; read the described methodology for look-ahead, survivorship, and selection red flags; check consistency across documents
What we doReproduction attempt of one claimed result in the manager's environment; timing and leakage checks; validation-design review; up to five principal findings
What we doFull review against the public review standard: methodology, data lineage, validation design, costs and capacity, parameter stability, version control; complete issue register
Deliverable
DeliverableFindings note (2–4 pp) with one-page summary, question list for the manager, and a 45-minute debrief
DeliverableFindings note (3–6 pp) with one-page summary, evidence from the reproduction attempt, and a 60-minute debrief
DeliverableFindings memo, issue register with severity and decision impact, reproducible evidence package, one-page summary, and debrief
Timeline
TimelineFive business days from complete materials and payment, plus up to three business days for the manager's factual check unless you waive it
TimelineFive to ten business days after environment acceptance and payment
TimelineTypically two to six weeks, agreed in writing
Fee
Fee$2,900 fixed, paid before work starts · 50% credited toward a Preflight of the same manager within 60 days
Fee$4,500 fixed, invoiced after environment acceptance · $2,250 credited toward a full validation within 60 days
FeeFrom $15,000, scoped in writing · 50% on the signed scope, 50% on delivery
Limitation stated in the note
Limitation stated in the note"Based on documents supplied by the manager; no code was executed and reproducibility was not independently verified."
Limitation stated in the noteScoped to one model, one pinned state, one execution path
Limitation stated in the noteNone beyond the agreed scope
Monitoring and re-validation. Annual re-validation of a manager already reviewed, and a drift check when the manager ships a new model version, are scoped on request at a fraction of the initial fee.
Co-sourced on request. Your analyst can sit in the reproduction: the checks, the harness notes, and the findings stay with your team afterwards. Same fee, same timeline.
Why the technical layer is the one that gets skipped
Operational due diligence checks the plumbing. Investment due diligence checks the story and the people. The model in between — the code that turned data into the track record — is usually accepted on the manager's word, because nobody on the buy side has the time or the seat to run it. Three reasons that is no longer comfortable:
The manager's side of the file is now regulated. (opens in a new tab) Under the SEC Marketing Rule (Advisers Act Rule 206(4)-1, in force for advisers since November 2022), an adviser that shows hypothetical or backtested performance must keep policies and substantiation behind it. An independent reproduction is the natural document on the allocator's side of that file.
Backtest overfitting is not a fringe risk. (opens in a new tab) With enough trials, a strategy with no real edge can be tuned to an impressive in-sample Sharpe ratio, and standard reports do not say how many trials were run (Bailey, Borwein, López de Prado & Zhu, Notices of the AMS, 2014).
Returns alone already say a great deal. (opens in a new tab) The best-known feeder-fund fraud was flagged before its collapse by a returns-based analysis that could not reproduce the reported numbers. Documents-only reads can catch real problems; ours prints their limits on the first page.
The manager cooperates; you stay in control.
A technical read of someone else's model needs their cooperation. Managers raising capital routinely agree under NDA — it is the same request a larger allocator's quant team would make. When they will not share code, we say so and read the documents instead, with the limitation printed on the first page.
Brief and fit check. You send what you have and the decision date. Within 24 hours we tell you which engagement fits, what we will need from the manager, and whether the date is realistic. One reviewer, one engagement at a time — if we cannot meet the date, we say so before you commit.
NDAs. We sign a two-way NDA with you and, where the manager requires it, the manager's own NDA. Findings are delivered to you and to no one else; we do not resell manager reviews and do not maintain a report library.
Access. Desk Read: documents only, through you or directly from the manager. Preflight and Validation: the manager provides code at a pinned state, one data snapshot, and an environment with read-only access; work runs inside that environment and only findings artifacts leave it. Client and manager materials are never used to train third-party AI models.
Factual check. Before the note is final, the manager receives the factual findings, not our conclusions, for a check lasting three business days. Corrections of fact are incorporated; disagreements are recorded as the manager's position. You may waive this step. For a Desk Read, the window starts when we send those findings and adds up to three business days after the five-business-day review. If no response arrives by the end of that Desk Read window, we record that and finalize the note using the available materials. We confirm the final delivery date before payment.
Delivery and debrief. The one-page summary and findings note go to you, followed by a debrief call. Fee and scope were fixed before step 3 and do not change with what we found.
Where we have done prior work for the manager under review, we disclose it before scoping and, where it was material, recuse and recommend an independent reviewer.
What we most often find in a manager's backtest.
A feature joined on its report date rather than the date it was actually available — invisible in the equity curve, decisive in diligence.
The target, or a transform of it, leaking into the training features.
Train/test contamination through overlapping windows or preprocessing fitted on the full sample.
A validation design that cannot support the claim: random splits on time series, or tuning on the evaluation period.
An evaluation window chosen after the results were seen.
A headline result that does not regenerate from the pinned code and the stated data.
A headline Sharpe that does not survive its own track length, the number of variants tried, or the costs the strategy would actually pay at the proposed size.
Six checks in every Desk Read · Six review areas in the public standard · Up to five principal findings, never padded · One named reviewer from scoping to handoff · Desk Read: Five business days from complete materials and payment, plus up to three business days for the manager's factual check unless you waive it
Eight questions worth asking before you allocate
Each maps to a check we run. Ask them first; the answers are often the whole story.
Which exact code version and data snapshot produced the numbers in the deck — and can they be pinned?
How many variants were tried before this one, and where are the ones that did not work?
Which features are stamped with the date they became available, rather than the date they refer to?
How were the train, validation, and test periods chosen — and were any chosen after the results were seen?
What cost, slippage, and capacity assumptions sit under the backtest, and at what size do they break?
Where does live performance diverge from the backtest, and what explains each divergence?
What has changed in the model since the track record began, and who signed off on each change?
Who can rerun the headline result today, from raw data, without the author in the room?
If the answers are complete and consistent, a Desk Read will say so. If they are not, you already know where a Preflight should start.
One page for the committee. Evidence for the file.
Every engagement ends with the same first page: the decision the read supports, the principal findings ranked by decision impact, and what remains unresolved — written to be forwarded as is. Behind it sits the evidence: what was reproduced, what was not, which checks were run and how, and — for Preflight and Validation — the artifacts a second reviewer could rerun.
What the first page looks like
Decision this read supports. The allocation, resize, or re-underwriting on the table, in one sentence.
Scope and access. What was read, in whose environment, at which pinned version — or, for a Desk Read, the statement that no code was executed and reproducibility was not independently verified.
Principal findings. Up to five, each with severity, decision impact, and a pointer to the evidence page.
What remains unresolved. What could not be checked with the access we had, and what would resolve it.
Questions for the manager. In the manager's terms, ready to forward.
Reviewer. Named. One reviewer from scoping to handoff.
Read the Desk Read sample → SAMPLE-002: a documents-only review of synthetic materials. No manager strategy code was executed; this is not client work or a real track record.
StatGazer / Public sample work productRef. SAMPLE-001
Illustrative · synthetic data · not client work · not a real engagement · not a track record
Model & Backtest Audit — Findings Memo
A material part of the headline does not survive point-in-time reconstruction.
A deliberately flawed synthetic strategy is rebuilt on point-in-time inputs so the decision-maker can see what survives.
Reported versus reproduced metrics from SAMPLE-001 — fabricated synthetic data, not client work, not a real engagement, and not a track record
Metric
Reported
Reproduced
Sharpe (annualized)
1.8
0.7
Annual return
14%
6%
Max drawdown
−9%
−21%
Hit rate
58%
52%
ConfirmedHigh
Look-ahead in the signal pipeline
A fundamental feature was joined on report date rather than its as-available date.
A synthetic demonstration — not client work and not a track record. This example shows our evidence format from a full validation exercise on fabricated synthetic data. A Preflight findings note follows the same evidence standard at first-read depth; a Desk Read note follows it at documents-only depth and says so on its first page.
We validate models we did not build. If StatGazer designed or materially modified a model, we say so and recommend an independent reviewer for its validation. When an allocator engages us, we take no fee from the manager under review. The fee is fixed and does not depend on what we find; findings are reported as found, and there is no "pass" product. Remediation after findings is separate, optional work — and any later re-validation of a model we helped remediate is disclosed as such in the findings note.
We do not select managers, build portfolios, run allocation mandates, or keep a library of manager reports to resell. You engage and pay us; the manager under review pays nothing. We deliver the note to you, and to your consultant if you wish. A Desk Read reviews supplied materials without executing the manager's code and states on its first page that reproducibility was not independently verified. A Preflight or full validation includes a scoped attempt to reproduce the agreed results in the manager's environment, with the agreed code and access.
Fit and boundaries.
Good fit
A systematic, quantitative, or model-driven manager whose case rests on a backtest, a model, or a track record — including digital-asset strategies (see Crypto due diligence for on-chain and exchange-specific checks).
A decision with a date: committee approval, allocation, resize, or re-underwriting.
A desk that wants a written paper trail with evidence, not a verbal opinion.
Not a fit
You need operational due diligence — custody, valuation, service providers, compliance infrastructure, cyber. Keep your ODD provider; this read complements it and does not replace it.
You want the manager's documents summarised rather than tested. Summaries are what AI tools and questionnaire platforms already do; a technical read reproduces the claim.
You need a recommendation on whether to invest, a rating, or a score. We document evidence; the decision is yours.
You need a guarantee, a certification, or an attestation. We do not certify models or managers.
The manager cannot provide even documents. There is nothing to read.
Frequently asked questions
Does the manager have to agree?
For a Preflight or a full validation — yes: we need code at a pinned state and an environment with read-only access, and only the manager can provide them. Managers raising capital routinely agree under NDA; it is the same request a large allocator's quant team makes. If the manager declines, a Desk Read works from documents alone and states that limitation on its first page.
We already use an investment consultant or OCIO. Why a separate technical read?
A technical read adds an evidence check to your existing diligence. A Desk Read checks headline statistics, costs, selection and capacity claims in the supplied documents; it does not execute the manager's code. A Preflight or full validation includes a scoped reproduction attempt in the manager's environment, with the agreed code and access. You engage and pay us; the manager under review pays nothing. We deliver the note to you, and to your consultant if you wish.
Who actually does the work?
One reviewer. The founder scopes the engagement, runs the reproduction, writes the note, and takes the debrief — no juniors, no hand-offs. That limits capacity to one engagement at a time, which is why the date is confirmed before you commit.
Is this an AI tool?
No. Document-summarising tools tell you what the manager wrote; a technical read tests whether it holds, by reproducing the result. We use software to run checks and to keep the evidence reproducible; the judgement, the findings, and the signature are one person's.
Who signs the NDA?
We sign a two-way NDA with you and, where the manager requires it, the manager's own form. Findings are delivered to you only. We do not resell manager reviews and do not keep a report library.
Does the manager see the findings?
The manager receives the factual findings, not our conclusions, for a check lasting three business days before the note is final. Corrections of fact are incorporated; disagreements are recorded as the manager's position. Conclusions remain ours, and the note is delivered to you. You may waive the factual check. For a Desk Read, the window starts when we send those findings and adds up to three business days after the five-business-day review. If no response arrives by the end of that Desk Read window, we record that and finalize the note using the available materials. We confirm the final delivery date before payment.
Can our own analyst take part?
Yes, on request. In a co-sourced read your analyst joins the reproduction; the checks, the harness notes, and the findings stay with your team afterwards. Same fee, same timeline.
Is this investment advice?
No. We report what the evidence shows about the model and the claim. We do not recommend whether to invest, size a position, or select a manager, and nothing we deliver is investment, legal, or tax advice.
Is this operational due diligence?
No. ODD covers custody, valuation, service providers, compliance, and cyber. We cover the model and the claim: reproduction, data timing, validation design, costs and capacity. The two are complementary; most allocators run both.
Do you take any fee from managers?
Not on allocator-engaged work. If we have previously worked for the manager under review, we disclose it before scoping and, where it was material, recuse.
Can you meet our committee date?
A Desk Read takes five business days from complete materials and payment, plus up to three business days for the manager's factual check unless you waive it. A Preflight takes five to ten business days after the manager's environment is accepted and payment is received. One reviewer works one engagement at a time. We confirm the final delivery date before payment and say no if it cannot be met.
What if the result is clean?
Then you have the evidence that it was checked, and a written statement of what remained unresolved. The fee is fixed and does not depend on the outcome.
What if the manager's numbers are right but the strategy would not survive costs at our size?
That is a finding. Cost, liquidity, and capacity assumptions are checked against the proposed allocation, not the manager's backtest size.
How do we pay?
By Stripe invoice — card or bank transfer. A Desk Read is paid before work starts. A Preflight is invoiced after the manager's environment passes acceptance, and scheduling is confirmed on payment. A full validation is 50% on the signed scope and 50% on delivery. Institutions can be invoiced directly; W-9, certificate of insurance, and NDA are available on request.
Terms at a glance
Fixed fee, agreed in writing before work begins; it does not change with what we find.
No fee, referral, or other compensation from the manager under review.
Two-way NDA with you; the manager's NDA where required; findings delivered to you only.
Preflight and Validation run in the manager's environment with read-only access; only findings artifacts leave it.
Client and manager materials are never used to train third-party AI models; retention and deletion terms are fixed in the engagement letter.
Professional liability (errors & omissions) and cyber coverage in place, underwritten by Hiscox; certificate of insurance on request.
Prior work for the manager under review is disclosed before scoping; where material, we recuse.
Engagement letter, NDA, W-9, and certificate of insurance are sent on request — usually the same day.
Founder · Direct delivery
Evgenii Azarov
The person who scopes your review is the person who runs it.
StatGazer is my one-person firm by design. I validate models for investment teams and review research statistics for people who publish — one reviewer across econometrics, statistics, machine learning, and financial engineering, with twenty years of private teaching and consulting behind every engagement.
20 years teaching & consulting, since 2006 · 100+ students from five countries · NDA-first, founder-only access · Insured — E&O and cyber (Hiscox)
Vendor onboarding in one email. W-9, certificate of insurance (professional liability and cyber, underwritten by Hiscox), NDA, and a one-page engagement letter — sent on request so your compliance desk can open the file the same day.
Use the narrowest engagement that fits the decision
StatGazer provides independent technical review engaged by the allocator. Manager technical due diligence is not operational due diligence, not a financial-statement audit, not a regulatory attestation, not a certification, and not investment, legal, or tax advice. Findings inform your own decision process; we do not recommend whether to invest.