Notes
Notes on grounding, evaluation and the parts that fail.
Short write-ups from the benchmark work: what a measurement actually showed, and what I changed because of it. First pieces are in progress.
No posts published yet
Until they're up, the same material is in the repositories: BENCHMARKS.md carries the methodology and full numbers, DESIGN.md carries the reasoning, and the known gaps are tracked as open issues.
Planned
Why citation grounding needs a second, independent pass
Designing an abstention rate you can report to auditors
What a synthetic corpus can and cannot tell you
Caption constraints as a dynamic program