GRPO reinforcement learning of LoRA adapters on Apple Silicon —
duel the base 35B against its RL-trained calibrated-honesty self,
live on the machine that trained it.
The machine thinking, in real time — prompts in, tokens out, tool
calls, prefill and decode timings. Whatever benchmark or agent is running
right now. Open only while a session is being shared, and shows a quiet
page otherwise.