Skip to content

Add named Excel table read and write connector - #1222

Open
mborodii-prog wants to merge 3 commits into
mainfrom
feature/1220-excel-table
Open

mborodii-prog wants to merge 3 commits into
mainfrom
feature/1220-excel-table

Conversation

@mborodii-prog

@mborodii-prog mborodii-prog commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Implements the Python portion of #1220; keep this PR draft until companion integration and live Excel verification pass.

Behavior

Add read: excel.table and write: excel.table for workbook-wide named tables. Reads require existing tables; writes can create missing destinations. WranglesXL supplies immutable __excel_tables snapshots through the existing recipe variables; Python reads headers and rows into a dataframe independently of selection. Writes emit ordered excel.table.write split payloads with target name and replace/append action. Replace is the default; later external selection batches append. Empty replacements are retained, and input snapshots never become memory write outputs.

Missing read tables error. Missing write tables are created on the configured worksheet and cell; existing workbook-wide tables retain their location. Names are case-insensitive. Headers are nonempty unique strings. XL requires matching existing column names and aligns order, resizes body rows, excludes totals on read, and rejects filtered/formula destinations and incoming formula strings in this first version. Schema is discovered by the existing generator; add user/transport documentation in docs/excel-tables.md.

This does not modify #943, selected-data connectors, saved recipes or deployment workflows. Existing Lambda variables and memory output transport carry the contract without a new top-level API field.

Validation

  • python -m pytest -c pytest-local.ini tests/connectors/test_excel.py -q: 70 passed, including actual recipe execution, composed sources, immutable input, ordered payloads, batch actions, invalid names/data/actions, empty results and generated schema validation.

  • Relevant offline regression run (excel, input, memory, matrix, recipe read/write): 154 passed using pytest-local.ini with WRANGLES_LOCAL_ONLY=1.

  • scripts/check_pytest_local_config.py: in sync (isolated Python environment).

  • git diff --check: passed.

Sheet/cell creation

  • write: excel.table accepts sheet and cell. Omitted cell becomes A1. Omitted sheet is the first 10 characters of <recipe_name>-<table_name> (fallback recipe name Recipe); invalid generated worksheet characters become underscores. Explicit worksheet names are preserved.
  • Python emits both resolved destination fields with each table output, including later external batches. Missing dataframe values are emitted as JSON null in table payloads without modifying the logical dataframe or read snapshots.
  • WranglesXL #1291 creates a missing worksheet/table, including a header-only table for empty output, after destination preflight. Existing tables stay in place. Occupied cells, table overlaps, colliding staged destinations and worksheet bounds overflow are rejected. The 10-character prefix can collide; use explicit distinct locations when needed.
  • Actual Python recipe tests cover union with differing columns, unmatched left join, concatenate and list reads, strict JSON serialization, explicit/default destinations and immutable snapshots. Real Python union/join output fixtures also pass the XL table creation tests.
  • Validation for this update: 154 Python offline regression tests; 83 tests across four XL Jest suites; targeted ES5 TypeScript check; both git diff --check runs. XL tests reuse the locally installed dependency tree with a scratch configuration and Windows sandbox realpath/TextEncoder setup; this is not a clean dependency-install or live Office test.
  • No merge, publication or deployment. Live desktop/web Excel + matching development-Lambda verification remains required, especially header-only creation, protection, surrounding cells and API failures.

Integration gates and limitations

  • Companion WranglesXL: wrangleworks/WranglesXL#1291.

  • WranglesJS table write-mode implementation is tested locally (16 tests plus targeted TypeScript compile); push denied HTTP 403 and fork denied organization policy. Maintainer write access is required to land the companion change before publishing.

  • Publish matching Python and generated schema, and validate a clean XL dependency build. No merge, publication or deployment performed.

  • Run live Excel/development-Lambda tests with synthetic data and recorded component versions, especially totals, filters, empty replacement, resize, surrounding cells, batching and cancellation.

  • Input snapshots must fit existing request limits; no table-input streaming. Top-level read/compositions only; hidden reads inside saved child recipes cannot dynamically fetch workbook tables.

  • Table outputs are staged before mutation, but Office writes are not transactional. Formula/filter rejection and unchanged column sets are documented compatibility limits.

  • Rollback: revert companion routing and connector changes; existing sheet/columns behavior remains available.

@mborodii-prog mborodii-prog self-assigned this Oct 9, 2026
@mborodii-prog
mborodii-prog marked this pull request as ready for review October 9, 2026 11:01

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Non-finite dataframe values can still produce invalid strict-JSON table payloads.

1 open finding
What changed in this PR

Adds named Excel table read/write support through WranglesXL’s existing variable and memory transports.

Changes:

  • Implements table snapshots, writes, validation, batching, and schema definitions.
  • Adds comprehensive connector tests and integration documentation.
  • Exposes the new documentation from the README.

Recommended disposition: Request changes

Next steps

  1. PR assignee: Define handling for infinite values, update the connector and regression tests, then run the focused local suite.
  2. AI agent: @codex address the non-finite JSON feedback, add ±infinity regression tests, and report the checks run
  3. Reviewer: Verify the fix, resolve the thread, and re-review after companion/live integration gates pass.
File Description
wrangles/​connectors/​excel.py Implements the named-table connector and validation.
tests/​connectors/​test_excel.py Tests reads, writes, schemas, batching, and destinations.
README.md Links the table documentation.
docs/​excel-tables.md Documents behavior and integration requirements.
.gitignore Allows the new documentation file.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

action = "append"
# A composed read may introduce NaN/pd.NA/NaT. Emit valid JSON cells
# without changing the dataframe returned to the recipe caller.
output = df.astype(object).where(_pd.notna(df), None)
@ebhills ebhills mentioned this pull request Oct 11, 2026
7 of 35 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants