tFirstnameMatch properties - 6.1

Talend Components Reference Guide

EnrichVersion
6.1
EnrichProdName
Talend Big Data
Talend Big Data Platform
Talend Data Fabric
Talend Data Integration
Talend Data Management Platform
Talend Data Services Platform
Talend ESB
Talend MDM Platform
Talend Open Studio for Big Data
Talend Open Studio for Data Integration
Talend Open Studio for Data Quality
Talend Open Studio for ESB
Talend Open Studio for MDM
Talend Real-Time Big Data Platform
EnrichPlatform
Talend Studio
task
Data Governance
Data Quality and Preparation
Design and Development

Component family

Data Quality

 

Function

tFirstnameMatch compares the first name column from the input flow with first names in an embedded reference index and outputs the matching first names.

This index has first names for about 162 countries, and it has more than 1000 reference first names for some countries. For further information, see About the reference index embedded in tFirstnameMatch.

Purpose

Helps ensuring the data quality of first names against a reference index in order to standardize data.

Basic settings

Schema and Edit Schema

A schema is a row description, it defines the number of fields to be processed and passed on to the next component. The schema is either Built-in or stored remotely in the Repository.

Since version 5.6, both the Built-In mode and the Repository mode are available in any of the Talend solutions.

One read-only column, FIRSTNAMEMATCH is added to the output schema automatically.

 

 

Built-in: The schema will be created and stored locally for this component only. Related topic: see Talend Studio User Guide.

 

 

Repository: The schema already exists and is stored in the Repository, hence can be reused in various projects and job designs. Related topic: see Talend Studio User Guide.

 

First Names

Select the column that contains first names.

 

Use Gender

Optional parameter: select this check box and then from the list, select the column that contains the gender. This will optimize system performance and give more precise results.

Expected genders are M (masculine) and F (Feminine).

 

Use Country

Optional parameter: select this check box and then from the list, select the column that contains the country ISO 3166-1 alpha-3 codes. This will optimize system performance and give more precise results.

 

Fuzzy Search

Select this check box if you want to get the best match possible, including approximate matches.

Advanced settings

tStatCatcher Statistics

Select this check box to gather the processing metadata at the Job level as well as at each component level.

Global Variables

ERROR_MESSAGE: the error message generated by the component when an error occurs. This is an After variable and it returns a string. This variable functions only if the Die on error check box is cleared, if the component has this check box.

A Flow variable functions during the execution of a component while an After variable functions after the execution of the component.

To fill up a field or expression with a variable, press Ctrl + Space to access the variable list and choose the variable to use from it.

For further information about variables, see Talend Studio User Guide.

Usage

This component is not startable and it requires input and output components.

Limitation/prerequisite

The index used to standardize the first names is embedded in this component. For the time being, it is able to handle Latin names.