diff --git a/README.md b/README.md index 4ea2ad2..19260a7 100644 --- a/README.md +++ b/README.md @@ -127,6 +127,7 @@ Most "awesome" lists are link dumps. This one is **annotated and verified**: eve - **[Hidden Technical Debt: Agent Evaluation Infrastructure](https://leehanchung.github.io/blogs/2026/06/13/hidden-technical-debt-agent-evaluation-infra/)** — Han-Chung Lee — · *blog* — Control plane / data plane; the **five surfaces** (output, trace, memory, environment, mechanistic); the empty-tool-result hallucination. - **[The Three Pillars of AI Observability](https://www.braintrust.dev/blog/three-pillars-ai-observability)** — Braintrust — · *blog* — Dataset reconciliation (living datasets); traces / evals / annotation. - **[Agent Trajectory Evaluations](https://arize.com/docs/ax/evaluate/evaluators/trace-and-session-evals/trace-level-evaluations/agent-trajectory-evaluations)** — Arize (AX docs) — · *docs* — Grading the path, not just the answer. +- **[Precise Records, Unstable Meanings](https://doi.org/10.5281/zenodo.21652317)** — Rolando Bosch (Hermes Labs) — · *paper* — Examines when precise agent telemetry fails to justify downstream evaluation claims, with a public evidence dossier and verification tools. 🆕 - **[AI Agent Metrics: How Elite Teams Evaluate](https://galileo.ai/blog/ai-agent-metrics)** — Galileo — · *blog* — A concrete agent-metric taxonomy (action completion, tool selection, etc.). - **[OpenInference semantic conventions](https://github.com/Arize-ai/openinference/blob/main/spec/semantic_conventions.md)** — Arize — · *tool/repo* — An OTel-based agent trace schema (tool, args, observation, latency, cost). - **[LangSmith Evaluation / Trajectory evals](https://docs.langchain.com/langsmith/evaluation)** — LangChain — · · *docs*. @@ -573,4 +574,3 @@ To the extent possible under law, [BenchFlow](https://benchflow.ai) and contribu -