Planning a Successful PoC

A proof of concept answers a specific question: does Senzing resolve your data well enough, and integrate easily enough, to be worth adopting? Evaluations succeed or fail on decisions made before any data is loaded.

This page covers those decisions. It is about scope rather than mechanics: the rest of the documentation covers how to install, map, load, and analyze, and each section below links to it.

Define what success means

Decide what you are measuring before you start, and write it down. Evaluations commonly weigh some mix of result quality, ease of data integration, speed to a usable result, cost of entry, how well the SDK embeds into your architecture, and data privacy.

Most PoCs come down to the first two: how good are the results, and how much work was it to get them. Rank your criteria, because they trade off against each other.

Senzing does not provide professional services or consulting, but has partners who do. Your team leads the evaluation, and unlimited support is included with an evaluation license, including help mapping your data sources. The people who run the PoC build the in-house expertise you will need in production.

Bring Senzing into your AI assistant

The Senzing MCP server puts a Senzing expert inside your AI assistant: Claude Code, Claude Desktop, Cursor, or any MCP-compatible tool. Throughout the evaluation it can help map your data sources, generate SDK code, troubleshoot errors, and search the Senzing documentation. It works from pre-fetched documentation and never sees your data.

Rightsize before you start

Three questions determine whether a PoC is achievable. Answer them honestly at the outset, because they constrain each other.

What data do you need? Which sources demonstrate the matching you care about, who owns them, and how quickly can you get access?

What systems are available? What can you provision, and for how long? Check your record volume against the v4 system requirements before committing to a scale.

Who is doing the work? How much of their time is genuinely available, and do they have the data and platform skills the plan assumes? The team needs people who understand the outcomes you want from the data, know its schema, and have access to it for the length of the PoC.

A billion-record evaluation staffed by one part-time person on a laptop will not finish. Scale the data down, the resources up, or the timeline out, but do not leave the mismatch in place.

Select the right data

Data selection determines what a PoC can demonstrate. The goal is a set that exercises real matching, not a tidy sample.

  • Use at least two or three sources. Cross-source matching is where entity resolution earns its keep, and a single source cannot show it. The truth set is shaped this way for the same reason: subjects of interest, reference data that enriches them, and a watchlist to screen against.
  • Pick sources that show your business cases. Choose sources that reflect the matching scenarios you care about, such as matching claims to customers, and that share at least a couple of attributes, such as name and address.
  • Take vertical slices, not random samples. A random sample scatters related records and destroys the overlap you are trying to observe. Slice by geography, date range, last names starting with the same letter, or another dimension that keeps connected records together.
  • Keep the mess. Typos, partial records, stale addresses, and inconsistent formatting are what production looks like. Cleaning them first tests a system you will never run.
  • Send every feature you have. Senzing uses what it is given, and withholding features suppresses matches it would otherwise find. The feature reference ranks each feature by how much it contributes to resolution, which is also a guide to what is worth chasing down before the PoC starts.
Do not naively mask or hash identity features such as dates of birth, national IDs, emails, and phone numbers. These carry much of the resolving power, and hashing destroys the partial and fuzzy comparisons that depend on them. Senzing sends no data outside your environment, so masking is rarely necessary to satisfy a privacy requirement.

If you have known matching pairs, assemble them into a truth set before you begin; even 50 to 100 labeled pairs is useful. A truth set turns “these results look reasonable” into precision, recall, and F1 scores you can defend. Where a specific scenario matters to the business, mock up records that exercise it directly rather than hoping the sample happens to contain one.

Choose a deployment platform

Match the platform to your record volume and to the skills already on the team. A native install on Linux is the closest to most production deployments; containers suit teams already working that way; the cloud guides reach hardware you do not have to buy for an evaluation.

Confirm the v4 system requirements for your record count first, and see database setup for repository options.

Map your data

Mapping tells Senzing what each field in your sources means. It is usually the shortest step in a PoC and the one teams over-plan: the out-of-the-box configuration covers most scenarios, an initial mapping usually takes less than thirty minutes per data source, and Senzing support will help you map your sources as part of the evaluation.

One entity per record. Each record describes one entity, typically a person or an organization. Do not mix the two in one record; where a source holds both, the entity specification covers how to mark each record’s type and how to keep the link between them, such as a person’s employer .

Start rough, then refine. Begin with a rough mapping and refine it once you have seen results, rather than perfecting one against data you have not yet resolved. The entity specification covers every attribute Senzing recognizes and how to map it.

Evaluate the results

Loading the data is not the finish line. The EDA tools are how a PoC turns resolved entities into an answer, and they cover the questions most evaluations are actually asking:

  • Entity exploration : Why did Senzing make each match, or decline to make it? The how and why commands in sz_explorer show the principles behind each decision.
  • Deduplication : How much did each source compress, and which records merged?
  • Cross-source screening : Which entities appear in more than one source?
  • Ambiguous matches : Where could a record plausibly belong to more than one entity?
  • Accuracy measurement : What are the precision, recall, and F1 against a truth set?

Results can also be exported as CSV and joined back to your source records, or replicated into a warehouse to query alongside existing data.

Check in regularly

Schedule progress reviews rather than waiting for a final readout. Regular checks surface a mismatched scope, a data access problem, or a mapping misunderstanding while there is still time to correct them. An evaluation that goes quiet for a month usually returns with an ambiguous result.

Reach out to [email protected] at any point during an evaluation.

If you have any questions, contact Senzing Support. Support is 100% FREE!