Portable CLI for judging RAG and LLM benchmark runs across local, OpenAI-compatible, and cloud providers — a deterministic quick mode, a paraphrase-tolerant LLM-as-judge mode, and a full per-case audit trail for every verdict.
Reads Claude Code’s own session transcripts and turns them into tokens, cost, time, and per-skill behavior — which prompt or skill is quietly ballooning your context, and which one keeps asking you questions.
A top-like live terminal dashboard for monitoring LLM inference servers on NVIDIA DGX Spark.