llmclean (Q7376804)

From MaRDI portal

!

This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:

LLM-Assisted Data Cleaning with Multi-Provider Support
Language Label Description Also known as
default for all languages
No label defined
    English
    llmclean
    LLM-Assisted Data Cleaning with Multi-Provider Support

      Statements

      0 references
      0 references
      Detects and suggests fixes for semantic inconsistencies in data frames by calling large language models (LLMs) through a unified, provider-agnostic interface. Supported providers include 'OpenAI' ('GPT-4o', 'GPT-4o-mini') <https://platform.openai.com>, 'Anthropic' ('Claude') <https://www.anthropic.com>, 'Google' ('Gemini') <https://ai.google.dev>, 'Groq' (free-tier 'LLaMA' and 'Mixtral') <https://groq.com>, and local 'Ollama' models <https://ollama.com>. The package identifies issues that rule-based tools cannot detect: abbreviation variants, typographic errors, case inconsistencies, and malformed values. Results are returned as tidy data frames with column, row index, detected value, issue type, suggested fix, and confidence score. An offline fallback using statistical and fuzzy-matching methods is provided for use without any application programming interface (API) key. Interactive fix application with human review is supported via 'apply_fixes()'. Methods follow de Jonge and van der Loo (2013) <https://cran.r-project.org/doc/contrib/de_Jonge+van_der_Loo-Introduction_to_data_cleaning_with_R.pdf> and Chaudhuri et al. (2003) <doi:10.1145/872757.872796>.
      0 references
      9 June 2026
      0 references
      PACKAGES.rds
      9 July 2026
      0 references
      0.1.0
      22 April 2026
      0 references
      0.1.1
      9 June 2026
      0 references
      0 references
      Rajesh Kaushal
      0 references
      0 references
      0 references
      0 references
      0 references

      Identifiers

      0 references