v4.4.0 Detailed Release Notes

Senzing 4.4.0 resolves more entities, more accurately, with enhanced country-specific address matching, Japanese-Kanji names, new languages, and website matching, plus faster loading and easier deployment.

What’s new in 4.4.0

The headline of 4.4.0 is resolution quality: enhanced country-specific address matching that nearly doubles the match rate on a set of difficult country-specific addresses (with zero regressions), initial Japanese-Kanji person names, new Ukrainian and Persian coverage, website/URL matching, much stronger business-name matching (distinct firms that share a generic phrase no longer over-merge, and a typo inside a long business name still matches), and smarter identifier handling (the new TAX_ID_TYPE attribute stops different kinds of tax identifier blocking a match, and placeholder values are ignored).

Alongside: faster loading, up to ~2× throughput depending on your bottleneck and up to ~6× faster reloading of already-seen data, lighter database load (up to ~45% fewer transactions and about half the rows read), correct UTF-8 on SQL Server and MySQL, a selectable high-concurrency locking strategy, native command-line tools on every platform (no Python runtime) with a smaller deployment footprint, and SDKs unified on one native core.

Many matching improvements apply to already-loaded data only after reprocessing. See Migration & Action Required . For the changes between v3 and v4, see v4 What’s New and v4 Breaking Changes .

Matching & Resolution Quality

Resolve more, and more accurately

This is the core of the release. Scores come from Senzing’s regression test data, run on 4.3 and again on 4.4. Cards without a score show the resolution outcome instead.

Enhanced country-specific address matching

Senzing’s country-specific address matching is significantly enhanced for Singapore, Japan, the UK, Canada, and the US, The same real-world address written with different street text, script, or formatting is recognized as a match, nearly doubling the match rate on difficult country-specific addresses. Each card shows a real pair and its measured 4.3 → 4.4 result.

Improved Singapore 6-digit postal = the building
A
1 Bayfront Avenue #01-05, Singapore, 018971
B
#01-05/06 Marina Bay Sands, 1 Bayfront Ave, Singapore, 018971
4.3 · LIKELY / 71 4.4 · CLOSE / 95
Singapore’s 6-digit postal code (018971) identifies the exact building, so two records at that postal, one written as a street address, one as the building name (“Marina Bay Sands”) with a slightly different unit (#01-05 vs #01-05/06), are recognized as the same place in 4.4, where 4.3 scored them apart. A genuinely different unit at the same building is still kept apart.
Improved United States ZIP+4 road aliasing
A
12842 Berlin Turnpike, Lovettsville, VA 20180-2342
B
12842 Route 287, Lovettsville, VA 20180-2342
4.3 · PLAUSIBLE / 66 4.4 · CLOSE / 95
The same delivery point under two legitimate road names (a highway route vs. its named road) now matches.
Improved Japan cross-script chōme-ban-gō
A
〒150-0002 東京都渋谷区渋谷1丁目1番1号
B
1-1-1 Shibuya, Shibuya-ku, Tokyo 150-0002
4.3 · NO_CHANCE / 47 4.4 · CLOSE / 90
The same Japanese address written in kanji, Latin, or full-width characters now matches (1丁目1番1号1-1-1).
Improved United Kingdom postcode + house #
A
100 Oxford Street, London W1D 1LL
B
100 The West End, London W1D 1LL
4.3 · PLAUSIBLE / 66 4.4 · CLOSE / 90
Same house number and postcode; differing street text no longer blocks the match. A different house number at the same postcode still doesn’t match.
Improved Canada 6-char postal is authoritative
A
100 Front Street West, Toronto, ON M5J 2L7
B
100 Harbourfront Centre, Toronto, ON M5J 2L7
4.3 · UNLIKELY / 62 4.4 · CLOSE / 90
Postal code plus house number identifies the delivery point; road text is treated as the less reliable field.

Person names and culture

4.4 significantly improves international person-name matching. A new model reliably identifies each name’s culture and applies language-specific rules, unlocking Japanese-Kanji, Persian, and Ukrainian handling that couldn’t be safely turned on before, and strengthening cross-script and abbreviated-name matching.

New Initial Japanese-Kanji person names
A
佐藤 (kanji)
B
Sato (Latin)
4.3 · NO_CHANCE / 42 4.4 · CLOSE / 98
Japanese surnames written in kanji now match their Latin spelling. Given-name coverage is partial and improving; because kanji is shared with Chinese, some Japanese names are still read as Chinese.
New Persian names match across script
4.3 · PLAUSIBLE / 76 4.4 · CLOSE / 95

FA محمدرضا ↔ Mohammadrezaa (native script vs an English spelling, misspelled)

Persian names now match their English spellings far better, and stay robust to typos: here the misspelled English spelling scores the same as a clean one.
New Ukrainian names matched as Ukrainian
4.3 · NO_CHANCE / 58 4.4 · CLOSE / 98

UA Григорій ↔ Hryhorii (same name, Cyrillic vs its everyday Latin spelling)

4.4 reliably recognizes a name as Ukrainian, so its Cyrillic form matches its everyday Latin spelling, where 4.3 read it with generic East-Slavic spelling and scored it apart.
Improved Abbreviated and nickname names find each other
A
Mhd Antoun
B
Mohamed Antoun
4.3 · never brought together, no match 4.4 · found & related
4.4 expands the name keys used to bring records together, so abbreviated and alternate spellings (Mhd ↔ Mohamed / Muhammad / Mohammed) now find each other on the name alone, where 4.3 would only compare them if some other feature (for example a shared phone) had already brought the records together.
New Names land in the right culture, so language-specific handling can finally be used
Persian and Ukrainian name handling existed before, but couldn’t be safely switched on: the engine couldn’t reliably tell a Persian name from a generic Arabic one, or a Ukrainian name from a generic East-Slavic one. 4.4 differentiates reliably between name cultures that share a script (Han → Japanese vs Chinese, Arabic script → Persian vs generic Arabic, Cyrillic → Ukrainian vs generic East-Slavic), which drives the Kanji, Persian, and Ukrainian gains above.

Organization names

4.4 substantially improves organization and business-name matching: coverage of common generic phrases is expanded, a typo inside a long business name no longer breaks the match, abbreviated names match their fuller form, and new Chinese and Burmese organization-name coverage extends matching to more spellings.

Improved Distinct firms sharing a generic phrase stay separate
A
Eastern Regional Community Health Services Foundation Inc
B
Western Regional Community Health Services Foundation Inc (same filing address)
4.3 · CLOSE / 96 · merged into one entity 4.4 · LIKELY / 85 · related, not merged
“Regional Community Health Services Foundation” is recognized as one generic phrase, so the distinctive first word decides: two different organizations stay separate but related, where 4.3 merged them.
Improved A misspelled, abbreviated name still matches
A
Springfld State U
B
Springfield State University
4.3 · NO_CHANCE / 49 4.4 · CLOSE / 94
A typo inside a long business name, combined with abbreviation, still matches.
New Chinese organization names match across spacing
A
BAI CHENG SHI
B
BAICHENGSHI (spaced vs unspaced pinyin)
4.3 · NO_CHANCE / 0 4.4 · CLOSE / 90
Spaced and unspaced pinyin forms of the same organization name now match.
New Burmese organization names match their English form
A
Yangon Kawporayshin (Burmese romanization)
B
Rangoon Corporation
4.3 · PLAUSIBLE / 74 4.4 · CLOSE / 94
A Burmese romanized organization name now matches its English equivalent.

Identifiers

New New TAX_ID_TYPE attribute
A
DE118621043 (Germany, VAT / USt-IdNr)
B
27/622/50488 (Germany, Steuernummer)
4.3 · conflicting IDs blocked the match 4.4 · different types, records merge
Many countries issue more than one kind of tax identifier (Germany’s VAT number and its Steuernummer, Italy’s Partita IVA and Codice Fiscale, and others). In 4.3 two of these for the same organization were treated as an exclusive conflict that blocked the match. The TAX_ID feature now carries a compared TAX_ID_TYPE attribute, so different types no longer conflict and the records resolve on their other shared data.
Improved Placeholder / junk identifiers ignored
A
National ID = n.a.
B
National ID = n.a. (unrelated records)
4.3 · linked on the shared placeholder 4.4 · not linked
Placeholder values no longer act as a shared identifier, so unrelated records are not linked by them.

Websites

New Website / URL matching
A
HTTP://Example.com/path
B
example.com/path
4.3 · NO_CHANCE / 0 4.4 · SAME / 100

example.com/pricing vs example.com/careers: different pages on the same host stay distinct (0/100)

The same website written different ways now scores as a match: different capitalization or http/https, and a Unicode domain alongside its punycode form, which are literally the same domain (中国.cn = xn--fiqs8s.cn). Genuinely different hosts, subdomains, and pages stay distinct. Website is corroborating evidence, not a standalone match key.

Loading & Performance

Faster loads, lighter database

Improved Higher loading throughput
4.3 → 4.4 · up to ~2× (workload-dependent)
4.4 batches and coalesces datastore writes (collapsing redundant row updates), relieving database round-trip bottlenecks. In testing this has delivered up to roughly 2× loading throughput versus 4.3. The gain depends on your system’s specific bottleneck. Resolution results are identical.
Improved Lighter database load
4.3 → 4.4 · up to ~45% fewer transactions · about half the rows read
4.4 issues up to ~45% fewer database transactions per record, with less write and log volume. Read I/O drops sharply too: with the per-entity feature store enabled, roughly half as many rows are read per record.
Improved Much faster reloading
4.3 · reprocessed in full every time 4.4 · up to ~6× faster

≈ Reloading records already seen (address-heavy)

New Opt-in Far fewer database reads on large entities
4.3
Large and hub entities issue progressively more database reads as the dataset grows, heavy on database I/O even when throughput is stable.
4.4
Substantially reduces datastore reads for large and hub entities. Where database I/O is the bottleneck (common at large scale), that is a major reduction. Off by default; opt in for large datasets.
Improved More efficient very large entities
4.4 trims per-operation overhead when resolving very large “hub” entities, so they process more efficiently as they grow.
New Opt-in More robust data-contention handling (PostgreSQL & SQL Server)
4.3
All databases used one strategy to coordinate concurrent entity writes; under heavy loading this added latency and long worst-case stalls.
4.4
On PostgreSQL and SQL Server you can opt into a database-native locking strategy, so bulk loads finish faster and far more predictably under heavy concurrency.
4.3 · under high contention 4.4 · lower worst-case latency · better overall throughput
The default strategy is unchanged, so existing deployments behave exactly as before unless you opt in.

Command-line Tools & SDKs

Native tools, one shared core

New Native command-line tools on every platform
4.3
Command-line tools were Linux-only and mixed Python and C++, requiring a Python interpreter to run.
4.4
The full tool set is now native and available on Linux, macOS, and Windows, with no Python runtime, so the Linux quickstart no longer requires one. The config tool also absorbs the former configuration-upgrade utility, and a new database tool creates, upgrades, and analyzes the datastore schema across all supported databases and single or clustered datastores, for example, flagging schema drift such as a missing index.
Improved SDKs on one shared native core
The Java, C#, and Python SDKs now sit on a single shared native core rather than three separate glue layers, giving more consistent behavior and error handling across languages. The Python SDK now follows the same shape as Java and C#, a language-level API layer over that shared native core.
Changed Clearer “not found” errors
4.3
Some record operations where a record disappeared mid-operation returned a generic engine error.
4.4
They raise the specific not-found type (Python SzNotFoundError / Java SzNotFoundException / C# NotFound).
The three tools that changed in this release. See Upgrade Schema to v4 and Upgrade ER Config to v4 for how they are used:

v4.0 - v4.3 v4.4.0 and later
sz_dbupgrade sz_dbtool upgrade
sz_configupgrade sz_configtool (configuration upgrades folded in)

Databases & Connectivity

Resilient and correct

Improved Improved UTF-8 on SQL Server & MySQL
SQL Server now uses UTF-8 collation and MySQL uses full 4-byte UTF-8 (utf8mb4), so emoji and rare characters are stored and compared correctly.
Improved Fail fast on a missing schema or table
A missing datastore schema or table is reported immediately and clearly, rather than surfacing later as an unrelated failure.

Platform & Packaging

Smaller, simpler deployment

New Fewer files in the deployment
The installed footprint is smaller, with fewer files to distribute and manage.
New Native macOS and Windows quickstarts
New guides cover installing and running Senzing natively, with Homebrew on macOS and Scoop on Windows .

Reliability & Operations

Steadier in production

Improved License misconfiguration fails fast
An invalid license string is rejected at startup rather than tolerated, so a misconfiguration surfaces immediately instead of during processing.
Improved Trace a stuck session to the exact thread
4.3
Every connection reported an identical client identity, so diagnosing a lock-holding session needed host/container access.
4.4
Each connection records its real process/thread identity in the database’s session metadata.

Fixed

Corrected in this release

Fixed Exclusive passport values no longer demoted
4.3
Records sharing an exclusive PASSPORT value were demoted to Possible Match when a junk passport value was also present.
4.4
The junk value is ignored, so the shared exclusive value resolves the records as it should.
Fixed Oversized native-script names no longer crash
Stack overflow on oversized native-script values during name comparison.

Migration & Action Required

Before you upgrade

Deployment and configuration notes for upgrading to 4.4.0 from v4.0.0 through v4.3.x.

Several of these require action outside Senzing itself. Review them against your deployment before upgrading.
Upgrading from v3 is a different path: the datastore schema and the configuration both have to be upgraded first. See Upgrade the datastore schema and Upgrade the configuration , then return to this list.
  • No schema change required when upgrading from any v4 version.
  • SQL Server: (see system requirements ) the datastore must use a UTF-8 database collation. On an existing datastore that is a rebuild/reload, not a live toggle, so plan for it in the upgrade window.
  • License: confirm your license string is valid: a corrupt value now fails startup instead of silently degrading to the evaluation license.
  • Configuration: some improvements are delivered through configuration and take effect only in a configuration that includes them: the WEBSITE feature, the TAX_ID_TYPE attribute, and the performance changes for very-common attributes. A new install’s default configuration already includes these; to add them to an existing configuration, contact Senzing Support for the steps.
  • Installers: if you are coming from 4.2.x or earlier, the Windows .zip and macOS .dmg were discontinued in 4.3.0. Install with Homebrew on macOS or Scoop on Windows , and update any scripts that pulled the old formats.
  • SDK exception handling: engine “record not found” conditions now raise the specific not-found type (SzNotFoundError / SzNotFoundException / C# NotFound). Adjust catch logic if you relied on the generic error.

Ready to upgrade

Review the migration checklist above, reprocess where new matching should apply to existing data, and reach out to your Senzing contact with any questions about this release.

If you have any questions, contact Senzing Support. Support is 100% FREE!