Lite SuiteLite Suite

Glass Box

Your local models, as a glass box.

See every token's confidence, what it almost said, and steer it. The token inspector turns raw logprobs into something you can actually read.

qwable-3.6-27b·replay · 37 tokens · entropy from top-5
<think> The user is asking about latency. Local models avoid network round-trips.</think> Local models run on your own GPU, so there is no network latency and the model stays warm in VRAM.
confidentuncertainhover a token for its top-5
Expert routing (MoE)instrumentation pending

Per-token expert activation is scaffolded — router instrumentation ships in a later pass.

Replay of a real inspection fixture — hover a token to open its top-5.

Per-token probability

Every token carries the probability the model assigned it — and the four runners-up it passed over. Click any token to see the full top-5.

Entropy heat

Confident tokens stay plain; uncertain ones glow gold. The heat is the model's own entropy, painted inline so you spot the wobble at a glance.

Prefill steering

Seed the assistant's own voice with a prefill and watch the distribution shift. Steer the model from inside its turn — not just its prompt.

Coming soon

MoE expert view — per-token expert routing is scaffolded in the inspector; router instrumentation ships in a later pass.

Run your own models with the lights on.

The Glass Box token inspector is built into LiteSuite — pull a GGUF, run it on your own GPU, and watch every token think.

Download LiteSuite

Windows 10/11 / Free 7-day trial