Data provenance
The record of where a given value came from and when it was last checked. Provenance attaches to the field rather than the row, so a single company record can have been assembled from several sources, each with its own capture time.
How it works
Every field carries its source, a capture timestamp, and a confidence score. Where sources disagree, the conflict is resolved programmatically and the resolution is written onto the field, so the next read of that record shows what changed and when.
What to evaluate
Ask a vendor to point at one field and trace it to the signal that produced it. Then ask what happens when that field is challenged and found wrong — whether the fix is made against the source, or patched over the top of it.
Identity resolution
The work of deciding that several records describe the same real-world entity. A company written four ways, or a person who changed jobs, has to collapse onto one identifier before any analysis of them means anything.
How it works
Records are resolved against a single identity graph, and the same canonical identifiers apply across every layer. A company found in search is the same company you enrich and the same one in the feed, rather than three records that contradict each other.
What to evaluate
Take a slice of your own data with known duplicates and ask what the vendor's graph makes of them. Ask how a record is handled when the evidence points both ways — a silent guess and a flagged ambiguity are very different answers.
Canonical record
The single surviving version of an entity after duplicates have been merged and conflicting fields resolved. It is the row everything else points at, and it is what stops the same company loading twice under two identifiers in the system you already run.
How it works
Matching runs against golden records, and the reason for any change is stored alongside the field rather than overwriting it quietly. A correction then lands in the feeds, in enrichment, and in search at once, because there is only one record to correct.
What to evaluate
Ask whether the canonical record is reachable from your side or only theirs. If you cannot read the current value, you cannot tell whether their correction reached you or whether you are still looking at your own stale copy.
Contact enrichment
Filling in missing attributes on a contact you already have — the current title, the company, the work email address — from what is known about that person and that company elsewhere.
How it works
Stale job titles and out-of-date company details are corrected against current signals rather than left to rot, and email addresses are checked for validity before the record is used rather than after a send bounces.
What to evaluate
Ask what an enrichment run costs when it finds nothing. Per-query billing charges you for the attempt; per-outcome billing charges you for the record that actually got fixed, which is the only one that changed anything.
Email verification
Establishing whether a work email address is deliverable and belongs to the person it is attached to, before it is used in an outbound program or written to a system of record.
How it works
Addresses are checked at the point the record is used, and the check travels with the record. A role-matched address and a guessed one are distinguishable, and the distinction is visible rather than implied.
What to evaluate
Ask what the provider does with an address that cannot be verified — dropped, flagged, or passed through. A pipeline that quietly discards records leaves you unable to audit what left your system.
Firmographics
Descriptive attributes of a company: what it does, which industry it sits in, where it operates, how large it is. Used to qualify accounts and to decide whether a company belongs in a segment at all.
How it works
Firmographic fields are re-checked on a rolling cycle, because they change more quietly than contact data. A reclassification, a new region, or a headcount band moving are corrections made as the cycle reaches them, not updates you have to request.
What to evaluate
Ask how a classification is decided when sources disagree, and whether the decision is recorded. An attribute you cannot trace is one you have to take on trust, and a segment built on it inherits that.
Data decay
The rate at which a dataset stops describing reality, because the world it describes has moved on. People change jobs, companies change shape, and addresses stop working — none of which is anyone's error at the moment of collection.
How it works
Re-verification runs on a rolling basis rather than in a batch before a sale. A record that is still accurate gets a fresh timestamp instead of being resold as new, so freshness is a property of the field and not a claim about the vendor.
What to evaluate
Ask what a still-accurate record looks like on its next delivery — a fresh timestamp, or silence. Freshness you can filter on yourself is worth more than a freshness claim you have to take on faith.
Deduplication
Removing records that describe the same entity, so a count of companies means a count of companies. It is a precondition for segmentation, attribution, and any per-record economics that assumes a record is a real thing.
How it works
Deduplication runs against the same identity graph as enrichment and search, so a merge decided once is consistent everywhere it is applied rather than re-litigated by each tool separately.
What to evaluate
Ask how the merge threshold is set and who sets it. Ask also what survives a merge — a wrong merge discards data quietly, and recovering it is harder than preventing it.
Natural-language search
Retrieval driven by a description of what you want rather than by a query language you have to construct. The request reads like a sentence and the system works out which records satisfy it.
How it works
Retrieval is semantic rather than keyword-matched, so intent carries weight that exact-string matching discards. Result sets are structured to feed straight into an outbound program or a product instead of needing to be reformatted first.
What to evaluate
Take a description you would give a colleague and ask to see the results. Then ask what a result that is wrong looks like — a narrow answer and a confidently wrong one are different problems to fix.
Ideal customer profile
A written definition of the companies a business is best suited to, expressed as attributes you can filter on — sector, size, geography, and the characteristics that actually predict a good customer rather than a plausible one.
How it works
Attributes that describe a company are kept current on a rolling cycle, so a profile written once does not silently drift out of date as the population it selects from changes underneath it.
What to evaluate
Ask whether the attributes a profile filters on are the same ones a buyer would answer on a call. A profile built from fields nobody would describe their own company in will exclude the accounts you meant to include.