Jump to content

Data cleansing: Difference between revisions

From Wikipedia, the free encyclopedia
Content deleted Content added
Line 13: Line 13:
* [[Data quality]]
* [[Data quality]]
* [[List of data quality vendors]]
* [[List of data quality vendors]]
* Automatic Data Cleaning and Formatting Software from QAS: http://www.qas.com/index.htm
* Data Cleansing Software http://www.winpure.com/
* Data Cleansing Software http://www.winpure.com/



Revision as of 19:48, 26 September 2006

Data cleansing is the act of detecting and correcting (or removing) corrupt or inaccurate records from a record set.

After cleansing, a data set will be consistent with other similar data sets in the system. The inconsistencies detected or removed may have been originally caused by different data dictionary definitions of similar entities in different stores, may have been caused by user entry errors, or may have been corrupted in transmission or storage.

Preprocessing the data will also guarantee that it is unambiguous, correct, and complete.

The actual process of data cleansing may involve removing typos or validating and correcting values against a known list of entities. The validation may be strict (such as rejecting any address that does not have a valid ZIP code) or fuzzy (such as correcting records that partially match existing, known records).

Data cleansing is synonymous with the less frequently-used term data scrubbing. Data cleansing differs from data validation in that validation almost invariably means data is rejected from the system at entry and is performed at entry time, rather than on batches of data.

See also

References

  • Han, J., Kamber, M. Data Mining: Concepts and Techniques, Morgan Kaufmann, 2001. ISBN 1-55860-489-8.
  • Kimball, R., Caserta, J. The Data Warehouse ETL Toolkit, Wiley and Sons, 2004. ISBN 0-7645-6757-8.