FundzData accuracy

Measured September 2026 · sample 2026-09b

How often is a Fundz signal on the right company?

Every funding round, hire, filing and acquisition we show is attached to a company. We check that attachment on a blind random sample, publish the result with its margin of error, and show you how to check it yourself. Most data vendors don’t publish this number.

Right company, exact legal entity
90.6%
95% CI 84.4%–96.9%
Right company family
(parent, subsidiary or sibling counts)
96.8%
95% CI 92.4%–100.0%

Weighted by how many links each data lane produces. Based on 873 decided links across 12 data lanes, drawn with seed e94c0ff9bc780226f1d53df8931de7a3. Excluding job postings, which are most of the volume: exact 89.6% (95% CI 86.7%–92.5%; n=797).

By data lane

Dot: exact legal entity. Ring: company family, shown where it differs. Bar: 95% confidence interval.

50%60%70%80%90%100%Form D filingsForm D filings: 100.0%100%SEC 8-K filingsSEC 8-K filings: 96.0%96%Executive hiresexecutive hires: 96.0%96%Funding newsfunding rounds (news): 93.5%94%Acquisitions: buyeracquisitions: acquirer: 93.5%94%Contractscontracts / agreements: 93.4%93%Job postingsjob postings (ATS): 90.8%91%WARN layoff noticesWARN layoff notices: 87.2%87%Office permitsoffice lease permits: 87.5%88%Product launchesproduct launches: 85.7%86%Trademarkstrademarks (USPTO): 82.9%83%Acquisitions: targetacquisitions: acquired company: 82.7%83%
Data laneLinks in windowExact companyCompany family
Form D filings1,207100.0% (95% CI 95.2%–100.0%; 77 of 77)100.0% (95% CI 95.2%–100.0%; 77 of 77)
SEC 8-K filings84696.0% (95% CI 89.0%–98.6%; 73 of 76)97.4% (95% CI 90.9%–99.3%; 74 of 76)
Executive hires2,13296.0% (95% CI 88.9%–98.6%; 72 of 75)97.3% (95% CI 90.8%–99.3%; 73 of 75)
Funding news1,70693.5% (95% CI 85.7%–97.2%; 72 of 77)93.5% (95% CI 85.7%–97.2%; 72 of 77)
Acquisitions: buyer2,85193.5% (95% CI 85.7%–97.2%; 72 of 77)93.5% (95% CI 85.7%–97.2%; 72 of 77)
Contracts2,41893.4% (95% CI 85.5%–97.2%; 71 of 76)94.7% (95% CI 87.2%–97.9%; 72 of 76)
Job postings145,94890.8% (95% CI 82.2%–95.5%; 69 of 76)97.4% (95% CI 90.9%–99.3%; 74 of 76)
WARN layoff notices4087.2% (95% CI 73.3%–94.4%; 34 of 39)97.4% (95% CI 86.8%–99.6%; 38 of 39)
Office permits10787.5% (95% CI 77.9%–93.3%; 63 of 72)91.7% (95% CI 83.0%–96.1%; 66 of 72)
Product launches1,37785.7% (95% CI 76.2%–91.8%; 66 of 77)85.7% (95% CI 76.2%–91.8%; 66 of 77)
Trademarks5,12282.9% (95% CI 72.9%–89.7%; 63 of 76)92.1% (95% CI 83.8%–96.3%; 70 of 76)
Acquisitions: target2,76282.7% (95% CI 72.6%–89.6%; 62 of 75)84.0% (95% CI 74.1%–90.6%; 63 of 75)

Filings keyed by a government identifier (SEC CIK) are near-perfect. The misses cluster where a record names a company only by name: two different companies with the same name.

What the audit found, plainly

What we miss, and duplicates

QuestionResult
Unlinked WARN notices that should have matched a company19.2% (95% CI 8.5%–37.9%; 5 of 26)
Unlinked office permits that should have matched a company4.0% (95% CI 0.7%–19.5%; 1 of 25)
Unlinked trademarks that should have matched a company13.6% (95% CI 4.8%–33.3%; 3 of 22)
8-K filings with no matched company where one existed97.5% (95% CI 87.1%–99.6%; 39 of 40)
Measured on the parser-company path (sampler v1). Since 2026-09-27 the product resolves 8-K filings by CIK first, which this figure does not measure; it describes the gap that change closes.
New companies that duplicated an existing one when created7.1% (95% CI 2.5%–19.0%; 3 of 42)
Companies with a duplicate row (random sample)18.9% (95% CI 11.6%–29.3%; 14 of 74)
lower bound: only generator-surfaced candidates are checked
Candidate pairs sharing a normalised name that are really one company84.2% (95% CI 62.4%–94.5%; 16 of 19)
Candidate pairs sharing a website domain that are really one company64.7% (95% CI 41.3%–82.7%; 11 of 17)
Candidate pairs sharing a LinkedIn company page that are really one company70.0% (95% CI 48.1%–85.4%; 14 of 20)
Candidate pairs sharing a parser company id that are really one company100.0% (95% CI 83.9%–100.0%; 20 of 20)
Candidate pairs sharing a CIK after zero-padding that are really one company66.7% (95% CI 43.8%–83.7%; 12 of 18)

Investors

Blind audit from stored evidence: 79.4% (95% CI 63.2%–89.6%; 27 of 34) 43 of 77 sampled links could not be decided from stored evidence

Checked against each round’s source article (2026-09-28, 160 of 320 sampled links readable): 93.7% (95% CI 88.8%–96.6%; 149 of 159).

Only 160 of the 320 sampled links could be measured. Links whose round source is BusinessWire, FinSMEs or a Google News redirect were excluded (the page could not be fetched or decoded into article text), as were links with no stored source URL or an unreadable page. The measured half therefore over-represents rounds reported by sources that serve plain article text, and the excluded sources may have a different precision. It is a second, independent measurement of investor-link precision, not a replacement for the blind audit's p_investor_link stratum (which could decide only 16 of 30 from stored evidence).

How we measure

Sampling
Every month a stratified random sample is drawn from the links Fundz made in that month. The seed is recorded before any row is labelled, and the order of rows within each stratum is md5(seed | stratum | row id), so anyone with the seed can re-draw exactly the same rows. The evidence each rater sees is frozen at draw time with a SHA-256 checksum, so a label can be re-checked even after the underlying data changes.
What is sampled
Precision strata sample event-to-company links per data lane (funding news, Form D, executive hires, contracts, acquisitions by role, product launches, job postings, WARN notices, permits, 8-K filings, trademarks) and investor-to-round links. Completeness strata sample records that were NOT linked and companies that were newly created, and ask whether an existing company was missed. Pair strata sample the current output of each duplicate-candidate generator. A duplicate stratum samples companies uniformly at random and asks whether each has a duplicate row.
Labelling
Each row is labelled by a primary rater from the frozen evidence. A second rater, a language model (disclosed by model name), labels the same rows blind to the primary label. Agreement is reported per stratum as raw agreement and Cohen's kappa. Every disagreement is re-read and adjudicated with a written note; the adjudicated label is final. Labels are append-only.
Who labelled
The primary rater for the first samples is Claude (an AI model) working in supervised sessions, and the second rater is also a model. No human labeller has labelled these samples. For that reason every audited row, its evidence and both labels are published, and a buyer is invited to draw and label a fresh sample of their own with the published seed procedure. Across 1,384 rows the two raters agreed 93.4% of the time (Cohen’s kappa 0.8939).
Statistics
Proportions are published with 95% Wilson score intervals. The cross-lane figure is a population-weighted stratified estimate with a finite-population correction. Rows a rater could not decide are reported as unknown and excluded from the denominator, never counted as correct.
What runs automatically
An automated analyst runs daily checks over new links and the duplicate backlog. A check may change data without a person only when its precision on a frozen, versioned gold set is at least 0.97 over at least 30 labelled rows, the change is reversible, it touches no company on an active customer's watchlist or lead list, and it is within a per-run cap. Everything else goes to a human review queue.
Publishing rule
A metric is published only with its sample size, its 95% interval and its measurement month, and only when its sample is at least 30 decided rows. Anything else is shown as not yet measured, with the reason. Limitations are always published beside the numbers.

Limitations

Predictions

We freeze every prediction on the day it is made and grade it once its horizon has passed. None of the current model’s horizons has closed yet, so its calibration is not yet measured:

Check it yourself

  1. The numbers on this page are version-controlled JSON: api.fundz.net/v1/data_quality.
  2. The sample is reproducible. Seed e94c0ff9bc780226f1d53df8931de7a3, sampler v1, window 2026-09-01 to 2026-09-27. ORDER BY md5('<seed>|<stratum>|' || <row id>) within each stratum's population filter; first n rows.
  3. Every audited row, its frozen evidence and both labels are in the data-room CSV for this month. Evaluating Fundz? We will draw a fresh sample with a seed you choose and show you every row — data@fundz.net.