Methodology
How TopETFList processes ETF data
The site uses a scheduled, versioned data pipeline. Visitors receive the most recent dataset that passed validation; their browsers do not scrape upstream websites.
Data selection and precedence
Approved upstream datasets have defined specialties and field-level precedence. Material disagreements remain recorded as conflicts instead of being silently overwritten. Upstream identities, URLs, and operational lineage are restricted to the authenticated administration area.
Validation before publication
- Only approved HTTPS homepages and CSV redirect hosts are accepted.
- Downloads have timeout, redirect, content-type, and 5 MB response limits.
- Required columns, ticker uniqueness, numeric values, and row counts are validated.
- Every imported field retains internal lineage, import time, and verification status.
- Builds fail when generated checksums, counts, tickers, provenance, or manifests disagree.
Status labels
- Source-reported
- Passed automated validation but has not been independently checked against an issuer source.
- Primary-verified
- Reviewed by a named person against an approved primary source.
- Conflict
- Two valid sources materially disagree on a value.
- Stale
- The source has exceeded its expected refresh window.
Important limitations
Current upstream responses do not provide a dependable financial-data as-of date. TopETFList therefore does not present an import time as the market-data date. Yields, returns, grades, and price-decay labels are source-reported unless explicitly marked primary-verified. Calculators are educational estimates and do not predict future distributions or returns.
