PerfectBlend — custom reward system (replace the proxy) #1
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
PerfectBlend — custom reward system (replace the proxy)
Status: PLAN ONLY. Design to be walked through one stage at a time, slowly
(per Kade: "if its not too much work we can go through it one by one slowly").
1. Current state (proxy — weak)
utils/reward_score/perfectblend.py(built commitdce31a5):sourcehas math cues) → numeric compare via thegsm8k last-number extractor. Genuine correctness signal where applicable. GOOD.
against the single reference "gpt" turn). WEAK: gameable, only measures surface
overlap with one reference, no instruction-following / semantic fidelity.
In the SqueakOpt MTMM audit,
perfectblendis currently a proxy trait (nocontent-grader peer) → any
judgeechoing it is flaggedUNVERIFIED. A betterreward moves it toward having real signal.
2. Goal
Compose a custom reward from reusable verifiers, iterated incrementally. Not a
single big model — start cheap and high-leverage, add components only after each is
reviewed and tested.
3. Stages (do in order, one at a time)
Reuse
scoring/constraint_check.py(theconstraintmethod used by ifeval /the SqueakOpt
instructiontrait). Verify the response follows the prompt'sexplicit constraints (length bounds, required keywords, format, casing). Applies
to ALL sources. Cheap, high signal.
with a local embedding model. Captures meaning beyond token overlap.
Verify an embedding endpoint exists on the mesh (check llama-services /
monitoring for an embed serve port) before committing to this stage.
scoring/llm_judge.py+otter_den LLM with a per-source rubric, for domains where the reference is a
style/solution exemplar. NOTE: per SqueakOpt audit a
judgeis UNVERIFIED —keep it a calibration signal, never gold.
poem/summary → constraint + semantic; code → exec where available). Weights
configurable in the handler/reward module.
4. Reuse inventory (don't rebuild)
scoring/constraint_check.py—constraintmethod (instruction trait).utils/reward_score/math.py/gsm8k.py— numeric extractors (S1).scoring/llm_judge.py+ otter_den LLM — S4.scoring/static_verifier.py,scoring/mtmm_methods.py— as needed.data_handlers/perfectblend.py— already splits conversations + carriessource;the reward module just needs to branch on
sourceper S5.5. Verification
tests/fixtures/build_fixtures.py).tests/test_perfectblend.py(constraint pass/fail,embedding delta, judge sanity).
perfectblendgains acontent-grader peer where constraint/semantic applies (no longer pure-proxy).
6. Pace
Implement S2 first; review; then S3; etc. "one by one slowly." No stage lands
until its unit test is green and the audit still passes.
7. Cross-references
~/.todos/randopt-datasets-plan.md— handler built commitdce31a5.~/.todos/terminal-bench-live-networking.md— the live eval (needs this reward)is what validates the reward end-to-end.
data_handlers/perfectblend.py,utils/reward_score/perfectblend.py.scoring/constraint_check.py,scoring/llm_judge.py.