Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,7 @@ env.bak/
venv.bak/
docs/*
!docs/ai_configuration.md
!docs/spreadsheet_recipes.md
!docs/ai_answers.md
!docs/extract_ai_user_guide.md
!docs/github-app-deployment.md
Expand Down
96 changes: 96 additions & 0 deletions docs/spreadsheet_recipes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
# Spreadsheet recipe inputs and writes

`excel.selected_data` reads the selected data supplied by WranglesXL for the
current recipe batch. It wraps `grid.selected_data`, which provides the same
options and behavior for grid hosts. Both support the usual read filters such
as `columns`, `not_columns`, and `where`.

```yaml
read:
- join:
how: left
left_on: ID
right_on: ID
sources:
- excel.selected_data: {}
- file:
name: products.csv
```

Excel batches only the selected rows. A batch size of 2 over five rows sends
2, 2, and 1 rows, with the matching selection headings, batch number, batch
total, and row count. External-only reads execute once, independently of the
selection's size. A mixed recipe executes its authored external reads for each
selected-data batch; scope ID-based queries to that batch's IDs when necessary.
An unfiltered external source can repeat rows across union batches.

Selected data must be present and nonempty. A missing or empty selected-data
operand in a join, union, or concatenate raises an error, including when its
read filters remove every row. Legacy `input` remains supported; a standalone
generic Python `input` can still represent an empty dataframe. Within a nested
recipe wrangle, these reads use the current parent dataframe, including earlier
transformations. They do not reread the original UI selection. The `read:
recipe` connector does not inherit that dataframe.

## Spreadsheet write modes

| Mode | Connector | Behavior |
| --- | --- | --- |
| `columns` | `excel.columns` | Insert new/changed output columns alongside the selected rows. |
| `sheet` | `excel.sheet` | Write all requested columns to the configured sheet/range. |

The complete calculation dataframe remains available. In columns mode, XL
compares the returned values with that batch's input and omits unchanged input
columns from the write. If everything is unchanged, nothing is inserted. The
number of output rows must equal the number of selected input rows; use sheet
mode for row-expanding or row-reducing results. Validation happens before
columns or headers are inserted for the batch.

```yaml
write:
- excel.columns:
columns: [Description, Category]
- excel.sheet:
name: Complete result
action: overwrite
```

Each write respects its own column selection. Sheet mode retains unchanged
input columns and the existing `name`, `cell`, action, and formatting options.
Later API batches append to the same resolved target. `matrix` repeats its
configured child writes: a child `excel.sheet` still uses sheet mode. Matrix
is a separate connector and is not another spreadsheet write mode.

Columns appearing in later batches are appended in first-seen order, with
values aligned by column name. Their final order cannot be known in advance.
If an input column first changes in a later batch, XL fills its earlier output
cells from the original selected values; it also retains subsequent unchanged
values in that output column. It stores range coordinates rather than buffering
the complete selection.

## Compatibility and execution

- `dataframe` keeps its Python return-value and column-selection semantics.
In XL columns mode, the runner translates a top-level legacy `dataframe`
write to `excel.columns`, including when combined with sheet writes.
- Recipes using selected data default to columns mode. Existing external-only
top-level reads retain their implicit sheet output. Legacy `output: range`
and `output: sheet` select the default mode; explicit `excel.columns` and
`excel.sheet` payloads select their own modes.
- An explicit columns write from an external-only read uses the selection only
as an output target and checks row alignment. The external read executes once.
- Nested recipe wrangles suppress side-effecting writes, including both Excel
writers. Use `dataframe` to shape a nested recipe's returned dataframe.
- Normal list runs use the Production version when present, otherwise latest.
Development Editor runs use the loaded/current editor definition, including
historical versions. The same snapshot controls planning and every batch.

The shared UI planning and dataframe projection helpers live in WranglesJS's
`recipe-execution` entry point. XL requests a full response with
`drop_unmodified_cols=false` and performs unchanged-column suppression at its
write boundary. Other API callers retain their existing default behavior.

The Python package containing these connectors must be promoted to the recipe
Lambda used by XL. A Python merge alone does not update that runtime. Likewise,
XL requires a WranglesJS package containing `recipe-execution`; publish/install
that companion package before integrating the XL change.
8 changes: 5 additions & 3 deletions schema/generate_recipe_schema.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,6 @@
import json
import logging
import yaml
import requests
import jsonschema

# Get the parent directory of the current file
Expand Down Expand Up @@ -125,7 +124,10 @@ def getConnectorDocs(schema_wrangles, obj, path):
schema['write']['dataframe'] = yaml.safe_load(
"""
type: object
description: Define the dataframe that is returned from the recipe.run() function
description: >-
Define the dataframe returned from recipe.run(). Retains its generic
Python return-value behavior. In Excel this is also legacy columns-mode
syntax; prefer excel.columns for an explicit spreadsheet write mode.
properties: {}
"""
)
Expand Down Expand Up @@ -234,7 +236,7 @@ def getMethodDocs(schema_wrangles, obj, path):
recipe_schema['$defs']['wrangles']['items']['properties'] = schema['wrangles']

# Validate the generated schema
jsonschema.validate(recipe_schema, requests.get('http://json-schema.org/draft-07/schema#').json())
jsonschema.Draft7Validator.check_schema(recipe_schema)

# Write final schema
with open('schema.json', 'w') as f:
Expand Down
46 changes: 46 additions & 0 deletions tests/connectors/test_excel.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,52 @@ def test_default_write():
)


def test_columns_and_sheet_writes_keep_complete_logical_dataframe():
memory.clear()
try:
original = pd.DataFrame({'ID': [1, 2], 'Value': ['a', 'b']})
result = wrangles.recipe.run(
{'write': [
{'excel.columns': {'columns': ['Value']}},
{'excel.sheet': {'name': 'Complete result'}},
]}, dataframe=original,
)
pd.testing.assert_frame_equal(result, original)
payloads = list(memory.dataframes.values())
assert payloads[0]['connector'] == 'excel.columns.write'
assert payloads[0]['columns'] == ['Value']
assert payloads[0]['data'] == [['a'], ['b']]
assert payloads[1]['connector'] == 'excel.sheet.write'
assert payloads[1]['columns'] == ['ID', 'Value']
assert payloads[1]['data'] == [[1, 'a'], [2, 'b']]
finally:
memory.clear()


def test_nested_excel_columns_does_not_emit_side_effecting_output():
memory.clear()
try:
result = wrangles.recipe.run(
{'wrangles': [{'recipe': {
'write': [{'excel.columns': {'columns': ['Value']}}]
}}]}, dataframe=pd.DataFrame({'ID': [1], 'Value': ['a']}),
)
assert result.columns.tolist() == ['ID', 'Value']
assert not memory.dataframes
finally:
memory.clear()


def test_generic_dataframe_still_shapes_the_python_return_value():
memory.clear()
result = wrangles.recipe.run(
{'write': {'dataframe': {'columns': ['Value']}}},
dataframe=pd.DataFrame({'ID': [1], 'Value': ['a']}),
)
assert result.to_dict('list') == {'Value': ['a']}
assert not memory.dataframes


def test_recipe_wrangle_in_batch_writes_all_rows_to_excel_sheet():
"""
Test the WranglesXL output connector path when a recipe wrangle is used
Expand Down
118 changes: 118 additions & 0 deletions tests/connectors/test_input.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import wrangles
import pandas as pd
import pytest

def test_read():
"""
Expand Down Expand Up @@ -79,3 +80,120 @@ def test_read_union():
df["header1"].values.tolist() == ["a", "b", "c", "d"] and
len(df) == 4
)


@pytest.mark.parametrize('connector', ['grid.selected_data', 'excel.selected_data'])
def test_selected_data_filters_current_batch_without_mutating_input(connector):
batch = pd.DataFrame({'ID': [3, 4], 'Value': ['c', 'd']})
result = wrangles.recipe.run(
{'read': {connector: {'columns': ['Value'], 'where': 'ID = 4'}}},
dataframe=batch,
)
assert result.to_dict('list') == {'Value': ['d']}
assert batch.to_dict('list') == {'ID': [3, 4], 'Value': ['c', 'd']}


@pytest.mark.parametrize('connector', ['input', 'grid.selected_data', 'excel.selected_data'])
@pytest.mark.parametrize('composition', ['union', 'join', 'concatenate'])
@pytest.mark.parametrize('data', [None, pd.DataFrame(columns=['ID'])])
def test_composed_selected_data_rejects_missing_and_empty(connector, composition, data):
params = {'sources': [{connector: {}}, {'test': {'rows': 1, 'values': {'ID': 1}}}]}
if composition == 'join':
params['on'] = 'ID'
with pytest.raises(ValueError, match='missing or empty'):
wrangles.recipe.run({'read': {composition: params}}, dataframe=data)


@pytest.mark.parametrize('connector', ['input', 'grid.selected_data', 'excel.selected_data'])
def test_composed_selected_data_rejects_filter_removing_all_rows(connector):
with pytest.raises(ValueError, match='empty after filtering'):
wrangles.recipe.run(
{'read': {'union': {'sources': [
{connector: {'where': 'ID > 10'}},
{'test': {'rows': 1, 'values': {'ID': 2}}},
]}}}, dataframe=pd.DataFrame({'ID': [1]}),
)


def test_generic_input_still_allows_an_empty_dataframe_outside_composition():
result = wrangles.recipe.run({'read': {'input': {}}}, dataframe=pd.DataFrame(columns=['ID']))
assert result.empty and result.columns.tolist() == ['ID']


@pytest.mark.parametrize('connector', ['grid.selected_data', 'excel.selected_data'])
@pytest.mark.parametrize('saved', [False, True])
def test_nested_selected_data_reads_transformed_parent_dataframe(connector, saved, monkeypatch):
child = {'read': {connector: {}}}
if saved:
monkeypatch.setattr(wrangles.data, 'model', lambda *_: {
'purpose': 'recipe', 'name': 'Synthetic child',
})
monkeypatch.setattr(wrangles.data, 'model_content', lambda *_: {
'recipe': f'read:\n - {connector}: {{}}\n',
})
child = {'name': '12345678-abcd-abcd'}
result = wrangles.recipe.run(
{'wrangles': [
{'convert.case': {'input': 'Value', 'case': 'upper'}},
{'recipe': child},
]}, dataframe=pd.DataFrame({'ID': [3], 'Value': ['batch three']}),
)
assert result.to_dict('list') == {'ID': [3], 'Value': ['BATCH THREE']}


def test_empty_external_result_is_not_an_empty_selected_data_error():
result = wrangles.recipe.run({'read': {'test': {'rows': 0, 'values': {'ID': 1}}}})
assert result.empty


@pytest.mark.parametrize('connector', ['input', 'grid.selected_data', 'excel.selected_data'])
def test_mixed_read_uses_only_current_batch_ids(connector):
# Synthetic lookup: the recipe author scopes the external query to the
# current invocation's IDs. Excel does not batch or deduplicate its rows.
catalog = pd.DataFrame({'ID': range(1, 6), 'Product': list('abcde')})
requested_ids = []
outputs = []
for ids in ([1, 2], [3, 4], [5]):
def query_products(ids):
requested_ids.append(ids)
return catalog[catalog.ID.isin(ids)].copy()

outputs.append(wrangles.recipe.run(
{'read': {'join': {'on': 'ID', 'how': 'left', 'sources': [
{connector: {}},
{'custom.query_products': {'ids': '${batch_ids}'}},
]}}},
dataframe=pd.DataFrame({'ID': ids}),
variables={'batch_ids': ids},
functions=query_products,
))
assert requested_ids == [[1, 2], [3, 4], [5]]
assert pd.concat(outputs).to_dict('list') == catalog.to_dict('list')


def test_composed_read_allows_empty_external_source():
result = wrangles.recipe.run(
{'read': {'union': {'sources': [
{'input': {}}, {'test': {'rows': 0, 'values': {'ID': 1}}},
]}}}, dataframe=pd.DataFrame({'ID': [7]}),
)
assert result.to_dict('list') == {'ID': [7]}


def test_generated_schema_exposes_shared_and_excel_selected_data(tmp_path, monkeypatch):
import json
import runpy
from pathlib import Path
import jsonschema

root = Path(__file__).resolve().parents[2]
(tmp_path / 'recipe_base_schema.json').write_text(
(root / 'schema/recipe_base_schema.json').read_text(encoding='utf-8'),
encoding='utf-8',
)
monkeypatch.chdir(tmp_path)
runpy.run_path(str(root / 'schema/generate_recipe_schema.py'))
schema = json.loads((tmp_path / 'schema.json').read_text(encoding='utf-8'))
for connector in ['input', 'grid.selected_data', 'excel.selected_data']:
jsonschema.validate({'read': [{connector: {'columns': ['ID']}}]}, schema)
jsonschema.validate({'write': [{'excel.columns': {'columns': ['Result']}}]}, schema)
1 change: 1 addition & 0 deletions wrangles/connectors/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
from . import concurrent
from . import duckdb
from . import excel
from . import grid
from . import file
from . import http
from . import memory
Expand Down
36 changes: 35 additions & 1 deletion wrangles/connectors/excel.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@
"""
import pandas as _pd
from . import memory as _memory
from . import grid as _grid
import logging as _logging


Expand Down Expand Up @@ -119,6 +120,37 @@ def _append_rows_by_column(saved: dict, df: _pd.DataFrame) -> bool:
return True


class selected_data:
"""Excel-facing wrapper for the shared grid selection contract."""

def read(dataframe: _pd.DataFrame = None):
return _grid.selected_data.read(dataframe)

_schema = {"read": _grid.selected_data._schema["read"].replace(
"Read the selected data", "Read the selected Excel data"
)}


class columns:
"""Spreadsheet write mode: columns alongside the selected rows."""

def write(df: _pd.DataFrame):
# Keep every requested column here. The grid writer compares the full
# result with its batch input and suppresses unchanged columns only
# in the columns-mode projection.
_memory.write(df, connector="excel.columns.write", orient="split")

_schema = {"write": """
type: object
description: >-
Spreadsheet write mode: columns. Insert new or changed output columns
alongside the selected rows in Excel. The output must retain row alignment
with the selection. Unchanged input columns are not duplicated.
additionalProperties: false
properties: {}
"""}


class sheet():
_schema = {}

Expand Down Expand Up @@ -179,7 +211,9 @@ def write(

_schema["write"] = """
type: object
description: Write to an excel sheet
description: >-
Spreadsheet write mode: sheet. Write all requested columns to an
Excel sheet/range, including unchanged input columns.
additionalProperties: false
properties:
name:
Expand Down
Loading
Loading