Skip to content

FEAT: expose explicit data shape metadata in load_data_source - #417

Open
biru-codeastromer wants to merge 1 commit into
sktime:mainfrom
biru-codeastromer:feat/expose-data-shape-and-scitype-metadata
Open

biru-codeastromer wants to merge 1 commit into
sktime:mainfrom
biru-codeastromer:feat/expose-data-shape-and-scitype-metadata

Conversation

@biru-codeastromer

Copy link
Copy Markdown
Contributor

Summary

Closes #416.

This adds explicit agent-friendly shape metadata to load_data_source and load_data_source_async responses so MCP clients can reason about loaded data handles without guessing from raw column names and dtypes.

What changed

  • add a shared metadata builder in Executor
  • expose:
    • target_scitype
    • target_variates
    • has_exog
    • exog_variates
    • index_type
    • n_target_columns
    • n_exog_columns
  • apply the same metadata enrichment to both sync and async load paths
  • add focused tests for:
    • datetime-indexed target with multivariate exogenous data
    • range-indexed target without exogenous data

Why this matters

This makes loaded data self-describing for agents before they choose estimators or workflow branches. In particular, it reduces ambiguity around:

  • whether a handle is univariate or multivariate
  • whether exogenous variables are present
  • whether the time index is datetime-like or integer/range based

Validation

  • .venv/bin/ruff check src/sktime_mcp/runtime/executor.py tests/test_data_sources.py
  • .venv/bin/ruff format --check src/sktime_mcp/runtime/executor.py tests/test_data_sources.py
  • .venv/bin/pytest tests/test_data_sources.py -q
  • .venv/bin/pytest -q

@Shashankss1205

Copy link
Copy Markdown
Collaborator

The idea (self-describing handles) still fits the current design and #416 is open. Two things are needed: (1) rebase onto main; there is one small conflict in executor.py where main now normalises the stored handle to a PeriodIndex (#531): keep that block and add your metadata call after it; (2) compute target_scitype / n_target_columns / target_variates instead of hardcoding 'Series', 1, 'univariate'; inspect_data already uses sktime.datatypes.check_is_scitype, so please reuse that so panel/multivariate handles are reported correctly.

@Me-Priyank

Copy link
Copy Markdown
Contributor

Had a look, target_variates is derived from y.ndim, but base.py:77 always returns a Series so it's always 1, and target_scitype is hardcoded to "Series". For a MultiIndex panel that comes back as "Series", which may be worse than leaving the field out. inspect_data.py:66-73 already uses check_is_scitype, reusing that would cover panels.

Also index_type is read before the PeriodIndex normalisation at executor.py:1236-1240, so on the default path it reports "datetime" for a handle whose index is actually a PeriodIndex.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[ENH] expose explicit data shape metadata for agent reasoning

3 participants