PerfectBlend — custom reward system (replace the proxy) #1

Open
opened 2026-07-19 10:51:08 +02:00 by kade · 0 comments
Owner

PerfectBlend — custom reward system (replace the proxy)

Status: PLAN ONLY. Design to be walked through one stage at a time, slowly
(per Kade: "if its not too much work we can go through it one by one slowly").

1. Current state (proxy — weak)

utils/reward_score/perfectblend.py (built commit dce31a5):

  • math-detectable sources (source has math cues) → numeric compare via the
    gsm8k last-number extractor. Genuine correctness signal where applicable. GOOD.
  • everything else → tolerant reference match (token-overlap / ngram-containment
    against the single reference "gpt" turn). WEAK: gameable, only measures surface
    overlap with one reference, no instruction-following / semantic fidelity.

In the SqueakOpt MTMM audit, perfectblend is currently a proxy trait (no
content-grader peer) → any judge echoing it is flagged UNVERIFIED. A better
reward moves it toward having real signal.

2. Goal

Compose a custom reward from reusable verifiers, iterated incrementally. Not a
single big model — start cheap and high-leverage, add components only after each is
reviewed and tested.

3. Stages (do in order, one at a time)

  • S1 — Math numeric (DONE). Keep as-is.
  • S2 — Instruction-constraint check (next, highest-leverage cheap win).
    Reuse scoring/constraint_check.py (the constraint method used by ifeval /
    the SqueakOpt instruction trait). Verify the response follows the prompt's
    explicit constraints (length bounds, required keywords, format, casing). Applies
    to ALL sources. Cheap, high signal.
  • S3 — Semantic similarity (embedding-based). Compare response ↔ reference
    with a local embedding model. Captures meaning beyond token overlap.
    Verify an embedding endpoint exists on the mesh (check llama-services /
    monitoring for an embed serve port) before committing to this stage.
  • S4 — Optional LLM judge (sparingly). Reuse scoring/llm_judge.py +
    otter_den LLM with a per-source rubric, for domains where the reference is a
    style/solution exemplar. NOTE: per SqueakOpt audit a judge is UNVERIFIED —
    keep it a calibration signal, never gold.
  • S5 — Composite weighting. Combine signals per-source (math → numeric only;
    poem/summary → constraint + semantic; code → exec where available). Weights
    configurable in the handler/reward module.

4. Reuse inventory (don't rebuild)

  • scoring/constraint_check.py — constraint method (instruction trait).
  • utils/reward_score/math.py / gsm8k.py — numeric extractors (S1).
  • scoring/llm_judge.py + otter_den LLM — S4.
  • scoring/static_verifier.py, scoring/mtmm_methods.py — as needed.
  • data_handlers/perfectblend.py — already splits conversations + carries source;
    the reward module just needs to branch on source per S5.

5. Verification

  • Offline fixtures already build a perfectblend dataset (tests/fixtures/build_fixtures.py).
  • Add per-stage unit tests in tests/test_perfectblend.py (constraint pass/fail,
    embedding delta, judge sanity).
  • After S2+, re-run the SqueakOpt MTMM audit to confirm perfectblend gains a
    content-grader peer where constraint/semantic applies (no longer pure-proxy).

6. Pace

Implement S2 first; review; then S3; etc. "one by one slowly." No stage lands
until its unit test is green and the audit still passes.

7. Cross-references

  • ~/.todos/randopt-datasets-plan.md — handler built commit dce31a5.
  • ~/.todos/terminal-bench-live-networking.md — the live eval (needs this reward)
    is what validates the reward end-to-end.
  • data_handlers/perfectblend.py, utils/reward_score/perfectblend.py.
  • scoring/constraint_check.py, scoring/llm_judge.py.
  • Forgejo: this work is tracked as the PerfectBlend reward issue (cross-linked).
# PerfectBlend — custom reward system (replace the proxy) Status: PLAN ONLY. Design to be walked through **one stage at a time, slowly** (per Kade: "if its not too much work we can go through it one by one slowly"). ## 1. Current state (proxy — weak) `utils/reward_score/perfectblend.py` (built commit `dce31a5`): - **math-detectable sources** (`source` has math cues) → numeric compare via the gsm8k last-number extractor. Genuine correctness signal where applicable. GOOD. - **everything else** → *tolerant reference match* (token-overlap / ngram-containment against the single reference "gpt" turn). WEAK: gameable, only measures surface overlap with one reference, no instruction-following / semantic fidelity. In the SqueakOpt MTMM audit, `perfectblend` is currently a **proxy** trait (no content-grader peer) → any `judge` echoing it is flagged `UNVERIFIED`. A better reward moves it toward having real signal. ## 2. Goal Compose a **custom reward from reusable verifiers**, iterated incrementally. Not a single big model — start cheap and high-leverage, add components only after each is reviewed and tested. ## 3. Stages (do in order, one at a time) - [ ] **S1 — Math numeric (DONE).** Keep as-is. - [ ] **S2 — Instruction-constraint check (next, highest-leverage cheap win).** Reuse `scoring/constraint_check.py` (the `constraint` method used by ifeval / the SqueakOpt `instruction` trait). Verify the response follows the prompt's explicit constraints (length bounds, required keywords, format, casing). Applies to ALL sources. Cheap, high signal. - [ ] **S3 — Semantic similarity (embedding-based).** Compare response ↔ reference with a local embedding model. Captures meaning beyond token overlap. *Verify an embedding endpoint exists on the mesh* (check llama-services / monitoring for an embed serve port) before committing to this stage. - [ ] **S4 — Optional LLM judge (sparingly).** Reuse `scoring/llm_judge.py` + otter_den LLM with a per-source rubric, for domains where the reference is a style/solution exemplar. NOTE: per SqueakOpt audit a `judge` is UNVERIFIED — keep it a *calibration* signal, never gold. - [ ] **S5 — Composite weighting.** Combine signals per-source (math → numeric only; poem/summary → constraint + semantic; code → exec where available). Weights configurable in the handler/reward module. ## 4. Reuse inventory (don't rebuild) - `scoring/constraint_check.py` — `constraint` method (instruction trait). - `utils/reward_score/math.py` / `gsm8k.py` — numeric extractors (S1). - `scoring/llm_judge.py` + otter_den LLM — S4. - `scoring/static_verifier.py`, `scoring/mtmm_methods.py` — as needed. - `data_handlers/perfectblend.py` — already splits conversations + carries `source`; the reward module just needs to branch on `source` per S5. ## 5. Verification - Offline fixtures already build a perfectblend dataset (`tests/fixtures/build_fixtures.py`). - Add per-stage unit tests in `tests/test_perfectblend.py` (constraint pass/fail, embedding delta, judge sanity). - After S2+, re-run the SqueakOpt MTMM audit to confirm `perfectblend` gains a content-grader peer where constraint/semantic applies (no longer pure-proxy). ## 6. Pace Implement S2 first; review; then S3; etc. "one by one slowly." No stage lands until its unit test is green and the audit still passes. ## 7. Cross-references - `~/.todos/randopt-datasets-plan.md` — handler built commit `dce31a5`. - `~/.todos/terminal-bench-live-networking.md` — the live eval (needs this reward) is what validates the reward end-to-end. - `data_handlers/perfectblend.py`, `utils/reward_score/perfectblend.py`. - `scoring/constraint_check.py`, `scoring/llm_judge.py`. - Forgejo: this work is tracked as the PerfectBlend reward issue (cross-linked).
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kade/SqueakOpt#1
No description provided.