LittleLearner Experiments: Teaching Math Beyond K–5
I trained 16 LoRA adapters on K–5-constrained LittleLearner models and found that useful mathematical skill transfer changed with…
I trained 16 LoRA adapters on K–5-constrained LittleLearner models and found that useful mathematical skill transfer changed with…
I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4…
English risky-financial fine-tuning transferred coherent emergent misalignment into Hindi, Marathi, and Urdu, while matched…
My first image-generation research experiment: teaching LongCat-Image-Dev source-conditioned editing, finding global…
I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on…
I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live…
How I built TOKENMAN for AMD Developer Hackathon ACT II: a deterministic, local-model, and Fireworks routing agent that climbed…
A public football fan simulation where AI fan agents from different countries react to matches, rivalries, upsets, predictions…
My ICML 2026 AdaptFM effort to optimize Qwen3.5-4B for low-latency inference on a single NVIDIA A10G through 100+ runtime…
I trained a sparse autoencoder on Qwen3-4B-Base, labeled its learned features, and tested one by steering the model toward…
Took a model from random initialization through 2B tokens of FineWeb-Edu on a single H100 and watched it learn next-token…
I didn't come close to the Parameter Golf leaderboard, but I still had a lot of fun running scattered H100 experiments on Modal…
How I built CrisisOps for the OpenEnv hackathon finale, then trained a small model with GRPO using TRL, Unsloth, and Modal.
A full fine-tune of Qwen3-1.7B on Wordle using OpenEnv, with a reward curve that actually went up and shaped rewards that taught…
My second RL experiment while studying GRPO: moving to Prime Intellect Lab, testing smaller and larger setups, and finally…
My first ever RL experiment: RLVR on GSM8K using Hugging Face TRL, Qwen2.5 1.5B Instruct, and an NVIDIA H100 on Modal, with notes…
How I study LLMs by going deep on specific topics instead of starting from math.
A first small-scale AI research experiment: cold-start SFT on Nanbeige4-3B-Base using 2,160 distilled reasoning triplets, trained…