-
Notifications
You must be signed in to change notification settings - Fork 0
M6: review, judge, and the baseline comparison #12
Copy link
Copy link
Open
Labels
fleetSpends model calls through the fleetSpends model calls through the fleetmilestoneA milestone tracking issue, M0 through M10A milestone tracking issue, M0 through M10reviewRound-trip, judge, sampling, the human funnelRound-trip, judge, sampling, the human funnel
Milestone
Description
Activity
Metadata
Metadata
Assignees
Labels
fleetSpends model calls through the fleetSpends model calls through the fleetmilestoneA milestone tracking issue, M0 through M10A milestone tracking issue, M0 through M10reviewRound-trip, judge, sampling, the human funnelRound-trip, judge, sampling, the human funnel
Review, judge, and the baseline comparison. Spends model calls.
No deterministic check can tell a correct translation from a fluent wrong one. A Vietnamese sentence that says the opposite of the English, with every role span intact, passes all nine invariants. This milestone is the answer to that, and it is honest about being a sample.
Checklist
review.py: round-trip prompt, judge prompt, the three-valued verdict, stratified seeded samplingreview calibrateand its published ratesreports/review.mdwith everydiffersentry in full, grouped by term and by rule rather than by fileMACHINE/tree against tier 1 against the human strings, onP03andG02Exit
Calibration rates published. If the judge does not catch dropped negations at a high rate, this issue says so and the round-trip is reported as unreliable rather than quoted as a result. The three-way comparison table is in a comment here. Tier 1's glossary and prompt fixes are merged and tier 1 is re-run against them.