Procedure - 7.0

Fuzzy matching

author
Talend Documentation Team
EnrichVersion
7.0
EnrichProdName
Talend Big Data Platform
Talend Data Fabric
Talend Data Management Platform
Talend Data Services Platform
Talend MDM Platform
Talend Real-Time Big Data Platform
task
Data Governance > Third-party systems > Data Quality components > Matching components > Fuzzy matching components
Data Quality and Preparation > Third-party systems > Data Quality components > Matching components > Fuzzy matching components
Design and Development > Third-party systems > Data Quality components > Matching components > Fuzzy matching components
EnrichPlatform
Talend Studio

Procedure

  1. In the Component view of the tFuzzyMatch, change the minimum distance from 0 to 1. This excludes straight away the exact matches (which would show a distance of 0).
  2. Change also the maximum distance to 2. The output will provide all matching entries showing a discrepancy of 2 characters at most.
    No other changes are required.
  3. Make sure the Matching item separator is defined, as several references might be matching the main flow entry.
  4. Save the new Job and press F6 to run it.
    As the edit distance has been set to 2, some entries of the main flow match more than one reference entry.

Results

You can also use another method, the metaphone, to assess the distance between the main flow and the reference, which will be described in the next scenario.