Written from 59 named sources · Aug 17 · first result Philanthropic Giving Intelligence: Source Architecture, Resolution Pipeline, and Build-vs-Buy Blueprint Scope: United States individuals — high-net-worth prospects, major-gift candidates, and foundation trustees. Two delivery tracks are specified: a low-touch licensed API stack and a custom Foundation Ledger. Every architectural decision below is driven by one governing constraint: the statutory record does not contain what most people assume it contains. The Epistemic Reality of Giving Data There is no public ledger of an individual's charitable giving in the United States. This is not a coverage gap that better engineering closes; it is a deliberate statutory design, and any system that pretends otherwise will ship fabricated facts. Under § 6104(b), the IRS must make tax-exempt organizations' annual information returns available to the public — but it is not authorized to disclose the name or address of any contributor to any tax-exempt organization other than a private foundation or a § 527 political organization. Contributor names and addresses are redacted before public release, and there has never been public disclosure of contributor identities outside those two categories [54]. Rev. Proc. 2018-38 then relieved non-501(c)(3) filers of Form 990/990-EZ from reporting contributor names and addresses at all, while requiring them to retain the data in their books for IRS examination [50]. Current Schedule B instructions confirm the operational shape: Schedule B is open to public inspection for a Form 990-PF filer and for a § 527 political organization filing 990/990-EZ; for all other filers, names and addresses are not required to be made available, and non-501(c)(3)/non-527 filers enter "N/A" in the contributor-name column [52]. ProPublica's own 990 XML parser reflects this reality — Schedule B is "typically marked as 'restricted'" in the machine-readable corpus [33]. The architectural consequence is precise. An individual's giving becomes an observed statutory fact only through four channels: Channel Statutory artifact What is observable Gives through a controlled private foundation Form 990-PF grants-paid schedule + Schedule B Grants out (recipient, amount, purpose) and substantial contributors in, both public for 990-PF filers [52][54] Gives to a § 527 political organization Form 990/990-EZ Schedule B; Forms 8871/8872 Contributor name, address, amount [52][16] Contributes to a federal political committee above the itemization threshold FEC Schedule A Name, address, occupation, employer for individuals exceeding $200 in a calendar year [27] Gift is publicly announced News, honor rolls, $1M+ gift registries Amount and designation, as reported [3] Everything else — estimated net worth, capacity bands, propensity and affinity scores, "interest in climate" — is modeled or inferred output, not observation. Commercial platforms are explicit that their scoring blends capacity, affinity, and propensity into configurable composites [11][9], and match rates on typical nonprofit files run roughly 70–90% [4]. A wealth screen tells you who can give; it does not tell you who has given, and vendors that lead with net worth carry no giving-history or mission-alignment layer at all [9]. Design rule #1: every field in the ledger carries an evidence class — observed_statutory, observed_first_party, observed_published, modeled, inferred — and no downstream surface may render a modeled or inferred value in the same visual register as an observed one. The Three-Tier Source Catalog Tier 1 — Statutory and public APIs (authoritative, free, high-latency) Source Access shape Coverage and grain Freshness / latency Notes ProPublica Nonprofit Explorer API REST v2, GET only, JSON/JSONP; /search.json, /organizations/:ein.json; 25 results per page; parameters for state, NTEE major group, 501(c) subsection Organization profiles plus Filing objects (990, 990-EZ, 990-PF) with 40–120 financial fields per filing; PDF and XML links [20] Structural lag is severe and documented: summary data processed in 2012–2019 calendar years generally covers 2011–2018 fiscal years [20][34] Usage constitutes agreement to ProPublica's Data Terms of Use [20]. Best free entry point for EIN discovery and officer/trustee surfacing [9] IRS bulk disclosure datasets Direct downloads; Form 990 series as XML; Pub. 78, Auto-Revocation, 990-N, EO BMF, Forms 8871/8872 [16][32] Full filing bodies including lettered schedules; political organization disclosures Datasets updated monthly; recent postings include Form 990 series XML 2026-06-17, Pub. 78 and Auto-Revocation 2026-06-09, 990-N 2026-07-06 [32] The canonical source of truth. Schemas, data dictionaries, and annotated forms are published [32] IRSx / 990-xml-reader Python library + CLI; standardizes 2013+ schema years to a canonical model; JSON/CSV/TXT output; supports 990, 990-EZ, 990-PF and schedules A–O and R [33] Object-level access by IRS object_id; repeating groups (officers, trustees, key employees, grants) Depends on your ingest cadence Schedule B is typically restricted in the corpus [33]. Note the maintainers' warning that the IRS stopped posting XML to AWS, so retrieval must be re-pointed at IRS bulk locations [33] GivingTuesday 990 Data Infrastructure 990 Data Lake (AWS/CLI), curated data marts for 990/990-PF/990-EZ, and an MVP API queryable by EIN [35] Research-ready extracts across the full corpus since 2023 [35] Follows IRS release cadence Removes most of the parsing burden if you do not want to own XML normalization Candid Essentials API v4 REST search over Candid's core nonprofit dataset; query generator in docs; Taxonomy API for controlled vocabularies [48][7] ~1.6M nonprofits with cause area, location, financials; Premier adds financials, people, and IRS-compliance data [12] Continuously refreshed; also available as seven pre-packaged datasets or custom builds [49] Taxonomy API is the correct normalization spine for subject, population, support strategy, and geography [7] Candid Grants API Endpoints for transactions, funders, recipients, summary [14] Institutional funder→recipient grant flows with subject/population/geography facets [14] Per Candid release cadence The single best instrument for foundation-level giving-area analysis; GuideStar and Foundation Directory capabilities are now unified in Candid search, with core nonprofit profiles and Forms 990 remaining free and unauthenticated [46][45] OpenFEC API RESTful, full-text and field-specific; /schedules/schedule_a/ for itemized individual receipts; Swagger schema published; free key at 1,000 calls/hour and 100 results/page, upgradable to 7,200 calls/hour [17] Individual contributions with name, address, occupation, employer above $200 per calendar year [27] Data updated nightly; Schedule A also published as a weekly database dump on Sundays; e-filing endpoints hold only the most recent four months [17] Governed by the sale-and-use restriction — see §4. This is the highest-freshness statutory source and the most legally constrained SEC filings Public search by individual name or ticker; free [3] Insider and stockholder disclosures, corporate filings; a capacity signal, not a giving signal [3][9] Filing-driven Useful for liquidity-event detection and officer/director confirmation Tier 2 — Commercial intelligence (licensed assertions plus proprietary models) Vendor What it actually delivers Access Reported pricing Freshness Structural limitation DonorSearch (EverTrue) Philanthropic-history-first screening: prior giving, political contributions, real estate, employment; DonorSearch Ai, Enhanced CORE, ProspectView Online 2 [3][9]; API returns philanthropic history, net worth, connections, ask amount [21] REST API plus 40+ CRM integrations; CRM add-ons metered as "API calls" purchased in plans against your EIN [21][39] No published rate card; the pricing URL 404s. Third-party reported: ~$4,000–$12,000/yr typical, ~$15,000–$22,000/yr enterprise, ~$1–$2 per record for one-time screening; legacy partner packages $228–$2,148/yr; no public figures for DonorSearch Ai [38][1] Vendor claims 1B+ data points updated weekly from dozens of sources and >90% accuracy on top prospects — vendor-claimed, not audited [3][38] Giving history is derived from public records rather than live asset data; no live profile monitoring; volume-oriented output [9] iWave / Kindsight 44+ vetted sources aggregated into 360° profiles; VeriGift database of 200M+ charitable gift records; ZoomInfo, Dun & Bradstreet, DatabaseUSA, Refinitiv, Insider Filings, RiXtrema; live profiles with real-time alerts; relationship mapping [9][11] Subscription and pay-per-use; CRM integrations [9] Published tiers: Starter from $4,500 (1,500 subscription screens, 1 user, 750 live profiles), Professional from $5,800 (5,000 screens, 3 users, 2,500 live profiles), Premium quote-only (20,000 screens, 5 users, 10,000 live profiles, 1 CRM integration, 7,000 append credits) [36] Real-time alerts on donations, real-estate transactions, insider filings [9] Deliberately not a one-click list generator; assumes analyst workflow [9]. Note aggregator price drift — stale directories still list $4,150 against a live $4,500 [38] WealthEngine (Altrata) Deep wealth profiles, income/asset/lifestyle indicators, predictive wealth scores, custom modeling; ~500B data points, ~250M pre-scored U.S. profiles, 5M+ refreshed weekly [9][11] API and Salesforce integration [9] Enterprise; quote-only [9] Refresh cycles reported as slower than newer cloud-native platforms [9] Capacity over charitable behavior; limited visibility into engagement, giving history, mission alignment; CRM integration beyond Salesforce is limited [9] Windfall Household-level net worth via deterministic algorithms; 25+ attributes; liquidity events, real estate, career intelligence [9][11] HTTP/JSON API with sub-second responses for individual requests, plus CRM integrations and file automations for full-database enrichment [37] Quote-only Weekly refresh [37][11]; serves 700+ nonprofits [11] No giving history, no affinity or propensity scoring — must be paired [9] GivingDNA / Dataro / Virtuous Insights First-party behavioral modeling layered on your CRM: propensity, upgrade, lapse, next-best-action; Insights surfaces suggested gift potential and real-estate holdings with Zillow verification links and confidence scores [11][8][4] Native CRM or integration Quote-driven; GivingDNA supports re-screening as often as monthly [11] Continuous against your own file These predict behavior; they do not discover external facts [8][4] Tier 3 — Affinity and soft signals (high value, low durability, human-review mandatory) Signal class Named sources Handling rule Announced major gifts Million Dollar List (Indiana University Lilly Family School of Philanthropy, $1M+ publicly announced gifts, free); Chronicle of Philanthropy "Philanthropy 50" and free $1M+ gift database [3] observed_published; store citation URL and publication date News and event monitoring Insightful Philanthropy monitors 14,000+ news and information sources with per-prospect alerting [3]; iWave live-profile alerts [9] Event, not fact; expires unless corroborated Board and trustee tracking 990 officer/director/trustee rosters via ProPublica and Candid Premier [9][12]; Prospect Visual relationship mapping and data appending [3] Affiliation ≠ personal gift. Store as affiliation, never as gift Professional and firmographic LinkedIn (education, employers, professional connections) [3]; iWave-licensed ZoomInfo/D&B/Refinitiv [9]; person-data platforms such as People Data Labs — contract terms must be reviewed for nonprofit and profiling use before ingest License-bound; never re-derive into a resold product without explicit rights Real estate Zillow lookups and county assessor records [3][9] observed_public for ownership, modeled for value Employer-side capacity Double the Donation: 30,000+ companies with corporate giving programs covering 26M+ employees [4] Employer program eligibility, not personal capacity Grant-seeking context Instrumentl: 200,000 funders, 12,500+ active opportunities; $179/$299/$499 per month tiers [3] Institutional, not individual AI-native research DonorAtlas searches open web and proprietary databases and links facts back to sources [10] Research aid only; require citation and named human reviewer per fact End-to-End Resolution Pipeline and Internal Ledger Architecture Your own CRM is the spine, not a peer input. Actual gifts to your organization, gift dates and amounts, event attendance, volunteer activity, stated interests, and consent preferences are the only fully defensible facts you hold — external data enriches them. Stage 1 — Identity resolution. Deterministic-first, probabilistic-second. Build a private identity graph rather than renting one; the market's own diagnosis is that only 23% of organizations report fully interoperable systems and only 50% hold a private identity graph, and that over-reliance on a single identifier such as hashed email cannot support interoperability alone [58]. Enterprise identity platforms resolve at billions-of-records scale precisely because they use multi-signal redundancy [57][58]. Emit a match_confidence on every link and never silently merge. Two US-specific landmines: FEC files legally contain up to ten fictitious "salted" pseudonyms per report, submitted separately and excluded from the public record, expressly to detect illegal list use [27][30] — an identity-resolution pipeline that treats FEC names as ground truth will ingest deliberate poison; and FEC "mailing address" may be a work address or P.O. Box, not a residence [27]. Stage 2 — Statutory ingestion. Anchor on EIN, not on person. Resolve controlled private foundations and board affiliations first, then walk the 990-PF grants schedule outward to recipients and the 990-PF Schedule B inward to substantial contributors [52][54]. Normalize subject and population coding through Candid's Taxonomy API so that "education," "climate," and "racial equity" mean one thing across the entire warehouse [7]. Budget for structural lag: the summary corpus demonstrably trails fiscal years by multiple years [20][34], and IRS datasets refresh monthly [32]. Stage 3 — Commercial enrichment. Data-minimized submission: send the narrowest identifying payload the vendor's match requires, request only the field allowlist you have a stated purpose for, and set per-vendor TTLs. Windfall's sub-second single-record API supports just-in-time enrichment at the moment a record qualifies, which is materially cheaper and more privacy-defensible than periodic whole-file screening [37]. Reserve batch screening for a scheduled full-file pass — the common cadence being an annual full screen plus continuous screening of new records [4]. Stage 4 — Confidence-scored composite. Every fact is a row, not a column overwrite. Required attributes on PROVENANCE: source_system, source_url_or_object_id, retrieved_at, publication_date, evidence_class, field_confidence, reviewer_id, review_status, ttl_expires_at. Required attributes on USE_RESTRICTION: no_solicitation_use, no_commercial_use, no_resale, no_automated_profiling, internal_fundraising_only, contract_id. Restrictions propagate: any derived score computed from a restricted input inherits the restriction. This is the mechanism that makes §4 enforceable in code rather than in a policy PDF. Composite confidence bands. Score 0.90–1.00 for first-party transactions and statutory facts tied to a filing object ID with a strong identity match; 0.60–0.89 for licensed third-party gift assertions and published announcements; 0.30–0.59 for vendor capacity models and single-source affinity signals; below 0.30 for unreviewed open-web extraction. Gift-officer surfaces render band and evidence class inline. Never present a capacity estimate as a giving history, and never let a wealth estimate alone drive differential treatment. Legal and Regulatory Compliance Matrix Boundary Rule as it actually reads Engineering control FEC sale-and-use restriction Information copied from FEC reports may not be sold or used to solicit contributions or for any commercial purpose, except that a political committee's name and address may be used to solicit that committee. "Soliciting contributions" expressly includes charitable contributions [31][27][30]. In AO 2003-24 the Commission held a 501(c)(3) could not use FEC contributor information for direct-mail educational material, because it could lead to a later communication containing a solicitation [27]. Using FEC data to purge or verify a commercial list is prohibited because it increases the list's commercial value (AO 1985-16) [27]. Academic research use is permissible (AO 1986-25) [27]; media republication is permitted where the principal purpose is not solicitation or commercial use [31][27] Tag every FEC-derived field no_solicitation_use=true, no_commercial_use=true. Physically segregate FEC data from any table that feeds list generation, mail merge, suppression, append, or ask-amount output. Block joins at the query layer, not by convention. Retain FEC signals for analysis and capacity context only, and require legal sign-off on any new consumer. Filter known salting artifacts [27][30] IRS Schedule B / § 6104 Contributor names and addresses are non-public except for private foundations and § 527 organizations [54][52]. Non-501(c)(3)/non-527 filers report "N/A" [52] Never attempt to reconstruct redacted contributor identity. Hard-code the disclosure matrix into the ingest validator: a Schedule B contributor name is only accepted from a 990-PF filer or a § 527 filer [52] CCPA/CPRA and the Delete Act A "data broker" knowingly collects and sells to third parties the personal information of consumers with whom it has no direct relationship [44]. Registration is annual, January 1–31, with a 2026 fee of $6,000 plus payment processing [42][44]; failure carries $200 per day plus back fees and Agency expenses [41][43]. Since August 1, 2026, registered brokers must access the accessible deletion mechanism at least once every 45 days and process consumer deletion requests [44]. Each distinct legal entity must register separately — subsidiaries cannot shelter under a parent — and must list all trade names and functioning websites [43]. SB 361 expanded disclosures to sensitive categories and to whether data was shared with foreign actors, law enforcement, or GenAI developers [44][41]. Independent third-party audits begin January 1, 2028, with registration disclosure of audit status from 2029 [44][41] This is the sharpest build-vs-buy fork. An internal ledger enriching your own constituents (direct relationship) sits outside the broker definition; a product that sells enriched person records to third parties almost certainly does not. If you are building the latter, budget registration, DROP account, a 45-day deletion-ingest job, per-entity registration, and audit readiness as first-class system components, not compliance afterthoughts. Implement suppression as an inbound pipeline stage that survives re-enrichment — otherwise deleted records silently reappear on the next vendor refresh GDPR/UK GDPR exposure Out of primary scope for a US-focused build, but engaged the moment EU/UK data subjects enter the file Enforce a residency gate at intake. Route non-US subjects to a separate lawful-basis workflow with legitimate-interest assessment, purpose limitation, and no automated profiling by default. Do not let a US-only enrichment contract be exercised against EU records Provider terms ProPublica use constitutes agreement to its Data Terms of Use [20]. Candid access is licensed per API/dataset [49][48]. Commercial prospect-research licenses commonly permit internal screening while restricting resale, redistribution, and use outside fundraising; DonorSearch CRM add-ons are metered per API call against your EIN [39][21] Maintain a machine-readable contract registry keyed to contract_id on every USE_RESTRICTION. Add a pre-deploy check that fails any new data consumer touching fields whose contract lacks that purpose Data minimization and fairness Statutory redaction and vendor restrictions both cut toward less Field allowlist per purpose; no collection of Delete Act sensitive categories [41] absent explicit necessity; TTL-based expiry of soft signals; a documented correction path so individuals can contest an identity match or profile value; never treat inference as fact in donor-facing or officer-facing copy Recommended Starting Stack: Two Tracks Track A — Low-Touch API Stack For fundraising organizations with an established CRM, no in-house data engineering, and a need to be productive this quarter. Roughly the profile DonorSearch is positioned against — organizations in the ~$2M–$20M annual revenue band [1]. Anchor on first-party data. Clean the CRM, resolve duplicates, and establish consent and communication-preference fields before any enrichment. Screening data that lands in a separate system loses most of its value [4]. License one philanthropic-history-first screener. DonorSearch if prior giving is the core question; iWave/Kindsight if you need relationship mapping, live profiles, and analyst depth in one place [9]. Buy at the published or reported entry tier — iWave Starter at $4,500 or Professional at $5,800 [36]; DonorSearch inside the reported $4,000–$12,000 band [38] — and require DonorSearch Ai priced as a separate line item, since no public figures exist for it [38]. Consider pay-per-record for one-time projects. At a reported $1–$2 per record, a 1,000-record screen at ~$1,500 is far below any subscription floor, while a 3,000-record screen approaches the subscription floor with no ongoing access [38]. Add capacity precision only if needed. Windfall for weekly-refreshed household net worth, accepting that it carries no giving history or affinity scoring and must be paired [9][37][11]. Layer free Tier 1 lookups for verification. ProPublica Nonprofit Explorer for EINs, filings, and trustee rosters [20][9]; Candid for cause area, foundation profiles, and taxonomy normalization [48][7][46]; SEC and county assessor records for asset confirmation [3][9]. Adopt one governance artifact: a per-fact provenance stamp in the CRM (source, retrieval date, evidence class, reviewer) and a hard rule that FEC-derived signals never enter solicitation workflows [31][27]. Indicative first-year cost: $5,000–$15,000 in licenses plus internal analyst time. No engineering headcount. Track B — Custom Foundation Ledger For product builders, multi-entity institutions, research teams, and anyone whose output will be redistributed, resold, or fed to models. This is where you own identity resolution and provenance rather than renting them. Statutory ingestion layer. IRS bulk 990 XML as the source of truth, monthly cadence [32]; IRSx or the GivingTuesday marts and Data Lake for normalization rather than writing your own XML parser [33][35]; ProPublica API for interactive search and reconciliation [20]; Candid Grants API transactions/funders/recipients/summary endpoints for institutional flow analysis and the Taxonomy API as the controlled vocabulary [14][7]; OpenFEC nightly plus the weekly Sunday Schedule A dump into a quarantined partition [17]. Private identity graph. Multi-signal, deterministic-first, confidence-scored, versioned, non-destructive. Own it rather than depending on a single identifier or a single vendor's resolution [58][57]. Ledger and provenance store. The entity model in §3, with USE_RESTRICTION propagation enforced in the query layer. Commercial enrichment as pluggable adapters. Windfall's HTTP/JSON API for just-in-time capacity at qualification moments [37]; a screener adapter (DonorSearch API or iWave) for philanthropic history [21][9]; behavioral scoring (Dataro, GivingDNA, or Virtuous Insights) computed on first-party data you already hold [8][11]