Pattern 01
Data cleaning and preprocessing
- Anomaly and gap detection: Rules plus ML and LLM classifiers flag null spikes, outliers and schema drift before a load lands in Snowflake or Databricks — not after a dashboard goes wrong.
- Text normalization: Abbreviations, locales and free-text synonyms standardised against an agreed dictionary, with human review for anything the rules are not confident about.
- Code mapping: Supplier and vendor codes aligned to internal MDM keys — deterministic maps first, fuzzy and semantic matching only for what is left over.