Generates the matching model that is used by the tMatchPredict component to automatically predict the labels for the suspect pairs and groups records which match the label(s) set in the component properties.
For further information about the tMatchPairing and tMatchPredict components, see tMatchPairing and tMatchPredict respectively.
tMatchModel reads the sample of the suspect pairs outputted by tMatchPairing after you label each second element in a pair, analyzes the data using the Random Forest algorithm and generates a matching model.
You can use the sample suspect records labeled in a Grouping campaign defined on the Talend Data Stewardship server with tMatchModel.
For further information about Grouping campaigns, see Adding a Grouping campaign to identify duplicate pairs.
-
Spark 1.6: CDH5.7, CDH5.8, HDP2.4.0, HDP2.5.0, MapR5.2.0, EMR4.5.0, EMR4.6.0.
-
Spark 2.0: EMR5.0.0.
For more technologies supported by Talend, see Talend components.