Skip to content

v0.4.0

Latest

Choose a tag to compare

@junclemente junclemente released this 08 Apr 19:04
4574e3a

[0.4.0] – 2026-04-08

Added

  • jcds.eda.show_null_rows(df, threshold=0.0) — added threshold parameter to filter rows missing more than a given proportion of values (e.g. threshold=0.5 for rows missing >50%)
  • jcds.eda.show_null_cols(df, threshold=0.0) — added threshold parameter to filter columns missing more than a given proportion of values
  • jcds.eda.show_outlier_summary(df, threshold=1.5, sort=True) — returns a DataFrame of outlier counts and percentages per numeric column, sorted descending
  • jcds.eda.inspect_row() — inspects a single row, shows null count and warns if mostly null
  • jcds.charts.hist_kde() — histogram with KDE overlay for numerical columns; supports individual mode (grid=False) and grid mode (grid=True, default) with configurable ncols, figsize, grid_figsize, export_func, and export_prefix parameters
  • jcds.charts.outlier_boxplots() — updated to support individual mode (grid=False) and grid mode (grid=True, default); added orient parameter ("v" or "h"), export_func, and export_prefix parameters
  • jcds.charts.cat_barplots() — multi-column categorical bar chart; supports individual mode (grid=False) and grid mode (grid=True, default) with configurable ncols, orient, top_n, max_unique, figsize, grid_figsize, export_func, and export_prefix parameters
  • jcds.reports.outliers() — new report combining outlier summary table and boxplot grid; supports threshold, orient, export_func, and export_prefix parameters
  • jcds.reports.show_dtypes() — dtype report; full dataset overview when called alone, deep dive into a single column when column= is provided
  • jcds.reports.categorical_summary() — summary report for categorical variables with value counts table and bar chart per column; filterable by max_unique and top_n
  • jcds.transform.drop_row() — drops a single row by position or label
  • jcds.transform.standardize_column_names() — canonical replacement for clean_column_names(), standardizes column names to snake_case
  • jcds.transform.drop_columns() — canonical replacement for delete_columns(), drops one or more columns

Removed

  • Deleted commented-out dead code (eda_guide_markdown) from inspect.py
  • Removed stale examples/eda_workflow.ipynb

Deprecated

  • jcds.transform.clean_column_names() → use standardize_column_names() instead
  • jcds.transform.delete_columns() → use drop_columns() instead

Docs

  • Added docs/getting_started.md — hand-curated EDA workflow guide organized by phase
  • Updated docs/index.md — updated version refs, test count, and links to new Getting Started page
  • Updated mkdocs.yml — added Getting Started and Changelog to nav