Skip to content

Add reduced part-code match evidence to search result scoring - #1229

Draft
ebhills wants to merge 9 commits into
mainfrom
codex/code-match-details-pra
Draft

ebhills wants to merge 9 commits into
mainfrom
codex/code-match-details-pra

Conversation

@ebhills

@ebhills ebhills commented Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator

Linked issue

Refs #1227 — accepted target: PY 1.20.5.
Refs #980 for the broader PRA cleanup; this PR does not close that issue.

What changes

Add standalone diagnostic fields inside each scored-result dictionary returned by compute.score_search_results:

  • part_code_matches: records ordered as match_type, match_level, match_source, matched_code, input_code.
  • part_code_match_count: distinct candidate/type/source evidence before redundancy removal, including the MPN repeated in part_codes; not an occurrence count within one field.

Matching uses the existing alphanumeric normalization helper. Case-only differences are exact; separator-only differences are stripped; containment within a larger code is partial. Short inputs still require a complete match rather than incidental containment.

Keep zero or one strongest MPN record, always first. Remove Code evidence already covered by the matched MPN, including normalized input substrings; collapse repeated matched Code values across title/snippet/URL. Prefer exact, then stripped, then partial; longer inputs and title/snippet/URL break ties.

The recipe continues to return one dictionary column or two parallel dictionary/summary columns. Numeric scores, scoring explanations, filtering, and result ordering are unchanged.

Recovery and final review

The original code-match-details-pra branch was pushed but not merged. It inherited unrelated dev-container, line-ending, and Copilot instruction commits, with merge conflicts in configuration files. This PR starts from current main (1d5e05a2) and carries only the eight matching commits plus one review-hardening/documentation checkpoint. The original branch and unrelated local Akeneo commits are preserved.

Final review fixed two diagnostic boundary problems:

  • A multiword overlap such as 085-196-225 M10 inside 085-196-225 M100 must remain partial and retain the full matched code.
  • Sentence punctuation must not make an exact code partial or allow a stripped variant elsewhere in the same field to hide the stronger exact evidence.

Document the output contract in docs/search_result_scoring.md, link it from the README, and update the recipe schema description.

How it was verified

  • Focused credential-free tests: .venv/Scripts/python.exe -m pytest -c pytest-local.ini tests/recipes/wrangles/test_compute.py -q --tb=short — 33 passed.
  • Full credential-free suite using scripts/test-local.ps1 (clears live-service credentials): 3,019 passed, 6 skipped, 139 deselected on Windows / Python 3.13.1 (171.09 seconds).
  • scripts/check_pytest_local_config.py — passed; 67 test files, 8 ignores, 22 deselects checked.
  • schema/generate_recipe_schema.py — schema generation and Draft7 validation passed; generated JSON remains local and is not committed.
  • Offline comparison against origin/main with 10 synthetic cases — existing scoring, filtering, ordering, and metadata identical after excluding only the two new fields.
  • git diff --check and git merge-tree --write-tree origin/main HEAD — clean, no merge conflicts; zero commits behind refreshed origin/main.

Direct regression coverage includes case-insensitive URL matches, key order, MPN-first selection, cross-source Code reduction, MPN substring suppression, raw counts, duplicate inputs/occurrences, no matches, single/two-column recipe output, complete-code boundaries, sentence punctuation, and score/filter invariance.

GitHub Actions CI run 38099564354 is in progress. No live search-provider or WrangleWorks-service acceptance was performed locally.

Compatibility and risk

Additive dictionary fields only; existing recipes keep the same output mapping. Strict consumers that validate exact dictionary keys may need to allow the two new fields. Reduced evidence intentionally omits redundant records; use part_code_match_count only as the documented pre-reduction evidence count.

Diagnostic match levels can be more precise than the existing scorer's exact/partial labels; no scoring-math changes are included. Matching dirty or unrelated short inputs can still produce incidental evidence under the existing normalization rules. No new dependencies, credentials, provider calls, version bump, deployment, or release publication are included.

Safest rollback: revert this focused PR and remove use of the two additive fields. The release version bump and promotion stay with the 1.20.5 release workflow.

Ready-for-review checklist

  • One human delivery owner is assigned: ebhills
  • The linked issue and intended milestone are correct: Release Planning for PY 1.20.5 #1227 / v1.20, targeting PY 1.20.5
  • The branch is current with main and has no merge conflicts
  • Focused tests pass
  • New or changed behavior has direct test coverage
  • Documentation/schema/configuration is updated where applicable
  • The PR contains no unrelated changes
  • The PR description reflects the branch's current scope and latest validation
  • Required GitHub Actions checks pass
  • One primary reviewer is requested only when this PR is ready

Draft handoff: implementation and local validation are complete. Keep Draft while CI is pending and the repository's five-PR Ready queue is full. After checks pass and a queue slot opens, the assignee should mark Ready and request one primary reviewer. No merge is authorized by opening this PR.

See the pull request workflow.

capture all code match information for downstream validation logic
Tighten search result part-code matching by treating code punctuation as part of token boundaries and reducing duplicate matches to the strongest explanation per occurrence. This also updates compute tests to cover redundant match removal, partial matches inside larger codes, and preferring match quality over MPN provenance.
 - remove codes matches contained within the MPN match
gives indication of how strongly the part codes show up in the content
@ebhills ebhills added this to the v1.20 milestone Oct 11, 2026 — with ChatGPT Codex Connector
@ebhills ebhills self-assigned this Oct 11, 2026
@ebhills ebhills mentioned this pull request Oct 11, 2026
7 of 35 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant