- Python 53.9%
- Jupyter Notebook 39.1%
- Shell 6.9%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
A thin OpenAI-compatible /v1/embeddings client plus a deterministic FakeClient for offline tests. get_embed_client() returns None unless SQUEAKOPT_EMBED_URL is set, so S3 stays a no-op and the reward falls back to S2 and the tolerant match until an endpoint is actually configured. A reward term that silently contributes zero is worse than one that is absent, hence the explicit None. |
||
| analysis/one_step_gradient | ||
| baselines | ||
| core | ||
| data | ||
| data_handlers | ||
| diffusion | ||
| distillation | ||
| docker | ||
| scoring | ||
| scripts | ||
| simple_1D_signals_expts | ||
| tests | ||
| utils | ||
| .gitignore | ||
| CUSTOM_DATASET_GUIDE.md | ||
| mtmm_audit.py | ||
| randopt.py | ||
| README.md | ||
| requirements.txt | ||
| validate_qwen36_params.py | ||
SqueakOpt
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights — audited with a multitrait–multimethod experiment.
SqueakOpt is a fork of RandOpt by Yulu Gan and Phillip Isola (MIT). It keeps the original optimizer and codebase and audits its central claim rather than just using it.
| Original paper | arXiv:2603.12228 |
| Original repository | github.com/sunrainyg/RandOpt |
| Original project page | thickets.mit.edu |
| SqueakOpt audit paper (LaTeX) | /home/kade/research/squeakopt/squeakopt.tex |
| SqueakOpt audit paper (build) | cd ~/research/squeakopt && make → squeakopt.pdf |
Paper | Original Repo | Project Page | 1D Colab
What SqueakOpt is
RandOpt's claim is a thicket regime: random Gaussian perturbations θ' = θ + g·ε around a
pretrained weight vector land on a dense neighborhood of task-specialized experts, evidenced by a
high solution density δ(m) (fraction of perturbations that beat the base model) and a high
spectral discordance D (the task-improvement rankings of different seeds are nearly
orthogonal → diverse, non-overlapping specialists).
SqueakOpt asks one skeptical question: how do we know the discordance is about capability and not about the ruler? Two seeds can rank differently on GSM8K because one genuinely reasons better — or because one happens to format the answer the way the numeric parser likes. RandOpt's own finding that only ~19% of its GSM8K gain is reasoning (the rest is format/style) is this confound in miniature: orthogonality comes for free from an unreliable or format-sensitive measurement, and discordance alone cannot separate "expert in math" from "expert in the math scorer."
SqueakOpt borrows psychology's multitrait–multimethod (MTMM) design to settle it. Each perturbation seed becomes a subject, scored repeatedly across capability traits (math, code, instruction, agentic) and methods (exact-match, symbolic verification, execution, static verifier, LLM-judge), on both in-distribution and fresh transfer item families. The decisive, hard-to-game test is transfer: do the seeds you select by one ruler on one item family stay on top when re-measured under a different ruler on different items? If rankings evaporate when the ruler changes, the thicket was dense in scorer space, not capability space.
The full write-up lives in ~/research/squeakopt/: squeakopt.tex (abstract → conclusion),
squeakopt.bib, and a Makefile that regenerates six vector figures from
figures/expectations.py and compiles to PDF via pdflatex + bibtex. The paper currently
compiles end-to-end against illustrative expected values; swapping in a real mtmm_report.json
(finalize once an MTMM run exists) regenerates the figures and results automatically.
What changed in the code (~/src/SqueakOpt)
This fork extends RandOpt's optimizer with the MTMM scoring/audit layer and begins hardening the reward scorers:
- Generation cache (
utils/gen_cache.py): persists each seed's raw response text per(seed, item)before the handler collapses it to a single mean reward. This is the enabling change — it makes multi-method scoring cheap and offline (generate once, score many). Env-gated bySQUEAKOPT_GEN_CACHE(off by default). - MTMM scoring layer (
scoring/):mtmm_methods.py— method registry (exact/symbolic/exec/static/judge/constraint), reusingutils/reward_score/{gsm8k,math,mbpp}.static_verifier.py— pyright/compile (no execution) for the CODE trait; pairs withexecso a seed that wins execution but fails the static check is flagged as an execution-format expert, not a code expert.constraint_check.py— IFEval rule-based constraint checks (the INSTRUCTION trait): keywords, length, language, format, casing, start/end, punctuation.llm_judge.py— a blinded Gemma-4-31B judge (no seed id leaked; different model family from the perturbed Qwen to break style bias), served via llama.cpp; 16 pytest cases green.
mtmm_audit.py— scores all(trait × method × family)cells, computes convergent / discriminant / transfer validity and discordance-shift, bootstrap CIs, and emitsmtmm_report.json+ a per-trait verdict (CAPABILITY/SCORER_ARTIFACT/INCONCLUSIVE).- Dataset handlers:
ifeval.pyandhumanevalplus.py(registered) for genuine method diversity on instruction-following and code;perfectblend.py(instruction-mix blend) andterminalbench.pyalready present. - PerfectBlend reward (
utils/reward_score/perfectblend.py): the original is a proxy (numeric for math-detectable sources, otherwise tolerant reference-match). We are replacing it with staged, reusable verifiers — S1 math-numeric (done), S2 instruction-constraint check reusingconstraint_check.py(done; seetests/test_perfectblend.py), S3 embedding-based semantic similarity (in progress). - Engine / precision split (verified): the weight perturbation
θ' = θ + g·εwithg ≈ 1e-3–5e-3needs fp16 resolution (ULP ≈ 9.8e-4) to register, so the pop-search runs on vLLM + fp16 (unquantized) on RunPod A100; the judge and final serving use quantized llama.cpp because every MTMM scorer operates on cached text, not weights.
Original RandOpt usage
Requirements
Option 1: Python / Conda
(optional) conda activate your_env
pip install -r requirements.txt
Option 2: Docker
From the directory containing SqueakOpt/:
| Step | Command |
|---|---|
| Build | docker build -f SqueakOpt/docker/Dockerfile_vllm -t randopt-vllm:latest . |
| Run | docker run -it --gpus all randopt-vllm:latest bash |
| Run (with data) | docker run -it --gpus all -v /path/to/SqueakOpt/data:/workspace/data randopt-vllm:latest bash |
Note: the Docker image tag (
randopt-vllm) is unchanged from upstream; it builds the SqueakOpt fork sources.
Run SqueakOpt
Post-train on your own dataset
Please follow the instructions in CUSTOM_DATASET_GUIDE.md
Post-train on a standard dataset
First download the data here: data/README.md
Then, from the SqueakOpt directory:
| Mode | Command |
|---|---|
| Single node | sbatch scripts/single_node.sh |
| Multiple nodes | sbatch scripts/multiple_nodes.sh |
| Local (no Slurm) | bash scripts/local_run.sh |
Distill top-k models into a single model
Please follow the instructions in distillation/README.md.
Run Baselines
Please follow the instructions in baselines/README.md
Citation
RandOpt (original)
@misc{gan2026neuralthickets,
title={Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights},
author={Yulu Gan and Phillip Isola},
year={2026},
eprint={2603.12228},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2603.12228},
}
SqueakOpt (this audit)
In preparation. Source of truth:
/home/kade/research/squeakopt/squeakopt.tex(build withcd ~/research/squeakopt && make).