Methodology

How TopETFList processes ETF data

The site uses a scheduled, versioned data pipeline. Visitors receive the most recent dataset that passed validation; their browsers do not scrape upstream websites.

Data selection and precedence

Approved upstream datasets have defined specialties and field-level precedence. Material disagreements remain recorded as conflicts instead of being silently overwritten. Upstream identities, URLs, and operational lineage are restricted to the authenticated administration area.

Validation before publication

  • Only approved HTTPS homepages and CSV redirect hosts are accepted.
  • Downloads have timeout, redirect, content-type, and 5 MB response limits.
  • Required columns, ticker uniqueness, numeric values, and row counts are validated.
  • Every imported field retains internal lineage, import time, and verification status.
  • Builds fail when generated checksums, counts, tickers, provenance, or manifests disagree.

Status labels

Source-reported
Passed automated validation but has not been independently checked against an issuer source.
Primary-verified
Reviewed by a named person against an approved primary source.
Conflict
Two valid sources materially disagree on a value.
Stale
The source has exceeded its expected refresh window.

Important limitations

Current upstream responses do not provide a dependable financial-data as-of date. TopETFList therefore does not present an import time as the market-data date. Yields, returns, grades, and price-decay labels are source-reported unless explicitly marked primary-verified. Calculators are educational estimates and do not predict future distributions or returns.