tJapaneseNumberNormalize - Cloud - 8.0

Text standardization

Version
Cloud
8.0
Language
English
Product
Talend Big Data Platform
Talend Data Fabric
Talend Data Management Platform
Talend Data Services Platform
Talend MDM Platform
Talend Real-Time Big Data Platform
Module
Talend Studio
Content
Data Governance > Third-party systems > Data Quality components > Standardization components > Text standardization components
Data Quality and Preparation > Third-party systems > Data Quality components > Standardization components > Text standardization components
Design and Development > Third-party systems > Data Quality components > Standardization components > Text standardization components
Last publication date
2024-02-20

Normalizes Japanese numbers (kansūji) to regular Arabic numbers.

Japanese numbers are often written using a combination of kanji and Arabic numbers with punctuation marks. Normalizing Japanese numbers can make those numbers more easily searchable and improve matching accuracy.

For example, tJapaneseNumberNormalize normalizes 3.2千 to 3200. This enables you to match the Japanese number "3.2千" with its Arabic counterpart "3200".

In local mode, Apache Spark 2.4.0 and later versions are supported.

This component is not shipped with your Talend Studio by default. You need to install it using the Feature Manager. For more information, see Installing features using the Feature Manager.

For more technologies supported by Talend, see Talend components.