[REFACTOR] Refactored the core code - #695
marcvergees merged 4 commits into
Conversation
… reworked Template creation and LLM calls to increase output accuracy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c6603d9120
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| best = max(scores, default=0.0) | ||
| matched = remaining.pop(scores.index(best)) if best > 0 else None |
There was a problem hiding this comment.
Consume zero-scoring rows before counting extras
When a populated output row has zero similarity to its expected row, this leaves the output row in remaining while also recording the expected row as unmatched. calculate_accuracy then penalizes the same mistake twice—once as a zero-scoring expected row and again as an unsupported extra row. For example, one correct and one wholly incorrect row score 1 / (2 + 1) = 33% rather than 1 / 2 = 50%. Match an available row even when its best score is zero, or exclude rows corresponding to unmatched expectations from the extra-row penalty.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
We have to discuss this @vharkins1. Should we get rid of the benchmark workflow?
There was a problem hiding this comment.
should we generalize a little bit this prompt removing explicitly the fact saying "You're a FEMA ICS form-filling engine". What would it happen if we export this to another country?
|
|
||
| SIGNATURE_TYPE = 6 | ||
|
|
||
| TYPE_TO_SCHEMA = { |
| from accuracy import _is_blank, match_rows | ||
| from config import JUDGE_CONTEXT_TOKENS, JUDGE_HOST, JUDGE_MODEL, JUDGE_TIMEOUT_SECONDS | ||
|
|
||
| JUDGE_THRESHOLD = 0.99 |
There was a problem hiding this comment.
it's not clear to me if we're batching the whole bunch of cases sent to LLM or not, could you clarify me this?
692539e
into
development-approach-d-benchmark
Replaces the Approach C/D fill code and the old benchmark with a small form-filling core (app/services/form_filler/) and a flat benchmark that runs directly against it. The goal is a working core to build outward from: this PR wires in the fill path only.
The benchmark for this specific code (qwen2.5:1.5b, 56 ICS narratives): 75.99% average accuracy (52 scored, 4 skipped because the model call failed or its output was cut off).
How filling works now
app/services/form_filler/filler.py fill(pdf_path, narrative, out_path, model):
The app reaches it through FileManipulator.fill_form, so /forms/fill and the Celery fill_form_task both use the new core. Ollama settings come from app/core/config.py; OLLAMA_TIMEOUT is new (default 300s).
Main Changes:
Added
Benchmark (benchmark/, replaces the old one entirely; see benchmark/README.md)
Docker: use an Ollama outside the container
Removed
Small
Testing