An MTMM Audit of Neural Thickets... Are “Diverse Task Experts” Real Capability Specialists, or Scorer Artifacts?
  • Python 53.9%
  • Jupyter Notebook 39.1%
  • Shell 6.9%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Balazs Horvath 6ff78634f1 Add the S3 semantic reward signal, gated on an embedding endpoint
A thin OpenAI-compatible /v1/embeddings client plus a deterministic FakeClient
for offline tests. get_embed_client() returns None unless SQUEAKOPT_EMBED_URL
is set, so S3 stays a no-op and the reward falls back to S2 and the tolerant
match until an endpoint is actually configured. A reward term that silently
contributes zero is worse than one that is absent, hence the explicit None.
2026-10-04 11:32:14 +02:00
analysis/one_step_gradient Some modifications 2026-06-20 10:01:50 +02:00
baselines Some modifications 2026-06-20 10:01:50 +02:00
core RandOpt: fix Phase-1 host-RAM OOM (vestigial 54GB store_base_weights + Ray object store) 2026-07-18 10:18:14 +02:00
data initial commit, basic setup 2026-03-01 11:52:12 -05:00
data_handlers SqueakOpt: S2 instruction-constraint check in PerfectBlend reward 2026-07-19 11:44:58 +02:00
diffusion Some modifications 2026-06-20 10:01:50 +02:00
distillation Some modifications 2026-06-20 10:01:50 +02:00
docker initial commit, basic setup 2026-03-01 11:52:12 -05:00
scoring Add the S3 semantic reward signal, gated on an embedding endpoint 2026-10-04 11:32:14 +02:00
scripts RandOpt: fix Phase-1 host-RAM OOM (vestigial 54GB store_base_weights + Ray object store) 2026-07-18 10:18:14 +02:00
simple_1D_signals_expts Some modifications 2026-06-20 10:01:50 +02:00
tests SqueakOpt: S2 instruction-constraint check in PerfectBlend reward 2026-07-19 11:44:58 +02:00
utils SqueakOpt: S2 instruction-constraint check in PerfectBlend reward 2026-07-19 11:44:58 +02:00
.gitignore RandOpt: PerfectBlend handler + proxy reward scorer (offline-testable) 2026-07-19 10:18:41 +02:00
CUSTOM_DATASET_GUIDE.md add CUSTOM_DATASET_GUIDE.md and Colab demo; update README.md 2026-03-12 00:15:16 -04:00
mtmm_audit.py RandOpt: SqueakOpt MTMM audit + IFEval/HumanEval+ handlers (offline-testable) 2026-07-19 09:34:41 +02:00
randopt.py RandOpt: SqueakOpt MTMM audit + IFEval/HumanEval+ handlers (offline-testable) 2026-07-19 09:34:41 +02:00
README.md docs(SqueakOpt): rewrite README for the RandOpt-audit fork (MTMM) 2026-07-19 11:44:58 +02:00
requirements.txt initial commit, basic setup 2026-03-01 11:52:12 -05:00
validate_qwen36_params.py Some modifications 2026-06-20 10:01:50 +02:00

SqueakOpt

Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights — audited with a multitrait–multimethod experiment.

SqueakOpt is a fork of RandOpt by Yulu Gan and Phillip Isola (MIT). It keeps the original optimizer and codebase and audits its central claim rather than just using it.

Original paper arXiv:2603.12228
Original repository github.com/sunrainyg/RandOpt
Original project page thickets.mit.edu
SqueakOpt audit paper (LaTeX) /home/kade/research/squeakopt/squeakopt.tex
SqueakOpt audit paper (build) cd ~/research/squeakopt && make → squeakopt.pdf

Paper | Original Repo | Project Page | 1D Colab

What SqueakOpt is

RandOpt's claim is a thicket regime: random Gaussian perturbations θ' = θ + g·ε around a pretrained weight vector land on a dense neighborhood of task-specialized experts, evidenced by a high solution density δ(m) (fraction of perturbations that beat the base model) and a high spectral discordance D (the task-improvement rankings of different seeds are nearly orthogonal → diverse, non-overlapping specialists).

SqueakOpt asks one skeptical question: how do we know the discordance is about capability and not about the ruler? Two seeds can rank differently on GSM8K because one genuinely reasons better — or because one happens to format the answer the way the numeric parser likes. RandOpt's own finding that only ~19% of its GSM8K gain is reasoning (the rest is format/style) is this confound in miniature: orthogonality comes for free from an unreliable or format-sensitive measurement, and discordance alone cannot separate "expert in math" from "expert in the math scorer."

SqueakOpt borrows psychology's multitrait–multimethod (MTMM) design to settle it. Each perturbation seed becomes a subject, scored repeatedly across capability traits (math, code, instruction, agentic) and methods (exact-match, symbolic verification, execution, static verifier, LLM-judge), on both in-distribution and fresh transfer item families. The decisive, hard-to-game test is transfer: do the seeds you select by one ruler on one item family stay on top when re-measured under a different ruler on different items? If rankings evaporate when the ruler changes, the thicket was dense in scorer space, not capability space.

The full write-up lives in ~/research/squeakopt/: squeakopt.tex (abstract → conclusion), squeakopt.bib, and a Makefile that regenerates six vector figures from figures/expectations.py and compiles to PDF via pdflatex + bibtex. The paper currently compiles end-to-end against illustrative expected values; swapping in a real mtmm_report.json (finalize once an MTMM run exists) regenerates the figures and results automatically.

What changed in the code (~/src/SqueakOpt)

This fork extends RandOpt's optimizer with the MTMM scoring/audit layer and begins hardening the reward scorers:

  • Generation cache (utils/gen_cache.py): persists each seed's raw response text per (seed, item) before the handler collapses it to a single mean reward. This is the enabling change — it makes multi-method scoring cheap and offline (generate once, score many). Env-gated by SQUEAKOPT_GEN_CACHE (off by default).
  • MTMM scoring layer (scoring/):
    • mtmm_methods.py — method registry (exact / symbolic / exec / static / judge / constraint), reusing utils/reward_score/{gsm8k,math,mbpp}.
    • static_verifier.py — pyright/compile (no execution) for the CODE trait; pairs with exec so a seed that wins execution but fails the static check is flagged as an execution-format expert, not a code expert.
    • constraint_check.py — IFEval rule-based constraint checks (the INSTRUCTION trait): keywords, length, language, format, casing, start/end, punctuation.
    • llm_judge.py — a blinded Gemma-4-31B judge (no seed id leaked; different model family from the perturbed Qwen to break style bias), served via llama.cpp; 16 pytest cases green.
  • mtmm_audit.py — scores all (trait × method × family) cells, computes convergent / discriminant / transfer validity and discordance-shift, bootstrap CIs, and emits mtmm_report.json + a per-trait verdict (CAPABILITY / SCORER_ARTIFACT / INCONCLUSIVE).
  • Dataset handlers: ifeval.py and humanevalplus.py (registered) for genuine method diversity on instruction-following and code; perfectblend.py (instruction-mix blend) and terminalbench.py already present.
  • PerfectBlend reward (utils/reward_score/perfectblend.py): the original is a proxy (numeric for math-detectable sources, otherwise tolerant reference-match). We are replacing it with staged, reusable verifiers — S1 math-numeric (done), S2 instruction-constraint check reusing constraint_check.py (done; see tests/test_perfectblend.py), S3 embedding-based semantic similarity (in progress).
  • Engine / precision split (verified): the weight perturbation θ' = θ + g·ε with g ≈ 1e-3–5e-3 needs fp16 resolution (ULP ≈ 9.8e-4) to register, so the pop-search runs on vLLM + fp16 (unquantized) on RunPod A100; the judge and final serving use quantized llama.cpp because every MTMM scorer operates on cached text, not weights.

Original RandOpt usage

Requirements

Option 1: Python / Conda

(optional) conda activate your_env
pip install -r requirements.txt

Option 2: Docker

From the directory containing SqueakOpt/:

Step Command
Build docker build -f SqueakOpt/docker/Dockerfile_vllm -t randopt-vllm:latest .
Run docker run -it --gpus all randopt-vllm:latest bash
Run (with data) docker run -it --gpus all -v /path/to/SqueakOpt/data:/workspace/data randopt-vllm:latest bash

Note: the Docker image tag (randopt-vllm) is unchanged from upstream; it builds the SqueakOpt fork sources.

Run SqueakOpt

Post-train on your own dataset

Please follow the instructions in CUSTOM_DATASET_GUIDE.md

Post-train on a standard dataset

First download the data here: data/README.md

Then, from the SqueakOpt directory:

Mode Command
Single node sbatch scripts/single_node.sh
Multiple nodes sbatch scripts/multiple_nodes.sh
Local (no Slurm) bash scripts/local_run.sh

Distill top-k models into a single model

Please follow the instructions in distillation/README.md.

Run Baselines

Please follow the instructions in baselines/README.md

Citation

RandOpt (original)

@misc{gan2026neuralthickets,
      title={Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights},
      author={Yulu Gan and Phillip Isola},
      year={2026},
      eprint={2603.12228},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2603.12228},
}

SqueakOpt (this audit)

In preparation. Source of truth: /home/kade/research/squeakopt/squeakopt.tex (build with cd ~/research/squeakopt && make).