● OpenAI-compatible · one Rust binary · MIT

The LLM gateway that learns to skip the LLM.

pankhllm started as an internal fix for LLM latency in production. Our agents kept paying a large model for the same kinds of decisions: which tool, which skill, which report. pankhllm learns those recurring decisions from your traffic, makes the safe ones itself in 0.2 ms on a CPU, and leaves everything unfamiliar to the LLM you already use.

decided by pankhllm: 0 sent to the LLM: 0 LLM calls skipped: 0%
The learning loop

Watch it learn, then get out of the way.

Your LLM keeps answering while pankhllm watches. Every decision it sees becomes a label; every night it retrains; every morning more questions never reach the LLM.

01 · OBSERVE

Every decision is a label

Planner choices, clean tool calls and explicit skill picks are recorded. Only the question's shape is kept.

<code><date><num>
02 · TRAIN

Overnight, on a CPU

pankhllm mine retrains the models and picks a confidence threshold for your target precision.

~1 sheld-out split
03 · DECIDE

Only when it can prove it

It acts only when confident, the wording is familiar, and the question fills the operation's parameters exactly.

0.2 ms0 LLM calls
04 · FALL BACK

Unsure? Your LLM answers

Nothing breaks while it learns. That answer becomes the next label.

fail-openshadow mode
Drop it in

Change one base URL. Keep your agent.

LangChain, Semantic Kernel, Microsoft Agent Framework, the Vercel AI SDK, or any OpenAI client. Small pankhllm clients ship for Python, TypeScript, JavaScript, C#, Go, Java, Rust, Ruby, PHP, C, C++ and the shell.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:4000/v1", api_key="...")  # was: your LLM provider
r = client.chat.completions.create(model="auto", messages=[{"role": "user", "content": "who covers R-3?"}])
# r.model_extra["pankhllm"]["model"] == "decision:coverage_owner"   (2 ms, no LLM call)

# or the pankhllm client:  pip install pankhllm
from pankhllm import Client
a = Client("http://localhost:4000").ask("who covers R-3?", tags=["private"])
print(a.text, a.model)
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "http://localhost:4000/v1", apiKey: "..." });
const r = await client.chat.completions.create({ model: "auto", messages: [{ role: "user", content: q }] });

// or the pankhllm client:  npm install @pankhllm/client
import { PankhClient } from "@pankhllm/client";
const a = await new PankhClient("http://localhost:4000").ask(q, { tags: ["private"] });
// Node 18+, plain JavaScript, no package needed
const r = await fetch("http://localhost:4000/v1/chat/completions", {
  method: "POST", headers: { "content-type": "application/json" },
  body: JSON.stringify({ model: "auto", messages: [{ role: "user", content: "who covers R-3?" }] })
});
const v = await r.json();
console.log(v.choices[0].message.content, v.pankhllm.model);   // "decision:coverage_owner"

// with the official openai package (CommonJS)
const OpenAI = require("openai");
const client = new OpenAI({ baseURL: "http://localhost:4000/v1", apiKey: "..." });

// or the pankhllm client:  npm install @pankhllm/client
import { PankhClient } from "@pankhllm/client";
const a = await new PankhClient("http://localhost:4000").ask("who covers R-3?");
// Semantic Kernel / Microsoft Agent Framework: change the endpoint, keep your plugins and loop
var kernel = Kernel.CreateBuilder()
    .AddOpenAIChatCompletion("auto", new Uri("http://localhost:4000/v1"), apiKey: "...")
    .Build();

// or the pankhllm client:  dotnet add package Pankhllm
var pk = new Pankhllm.PankhClient("http://localhost:4000");
var a = await pk.AskAsync("who covers R-3?", new AskOptions { Tags = ["private"] });
Console.WriteLine($"{a.Text} via {a.Model}");
import "github.com/singhpratech/pankhllm/sdks/go"

pk := pankhllm.New("http://localhost:4000")
a, err := pk.Ask(ctx, "who covers R-3?", pankhllm.Options{Tags: []string{"private"}})
if err != nil { log.Fatal(err) }
fmt.Println(a.Text, a.Model)   // a.Model == "decision:coverage_owner"

// any OpenAI Go client works too: set the base URL to http://localhost:4000/v1
import dev.pankhllm.PankhClient;

var pk = new PankhClient("http://localhost:4000");
var a  = pk.ask("who covers R-3?", new PankhClient.Options().tags(List.of("private")));
System.out.println(a.text() + " via " + a.model());

// Spring AI / LangChain4j / the official OpenAI Java SDK: base URL http://localhost:4000/v1
use pankhllm_client::{Client, Options};

let pk = Client::new("http://localhost:4000");
let a = pk.ask("who covers R-3?", &Options { tags: vec!["private".into()], ..Default::default() }).await?;
println!("{} via {}", a.text, a.model);

// the router itself:  cargo install pankhllm  &&  pankhllm serve --port 4000
require "pankhllm"

pk = Pankhllm::Client.new("http://localhost:4000")
a  = pk.ask("who covers R-3?", tags: ["private"])
puts "#{a.text} via #{a.model}"

# ruby-openai works too:  OpenAI::Client.new(uri_base: "http://localhost:4000/v1")
use Pankhllm\Client;

$pk = new Client('http://localhost:4000');
$a  = $pk->ask('who covers R-3?', ['tags' => ['private']]);
echo $a['text'], ' via ', $a['model'];

// openai-php works too: OpenAI::factory()->withBaseUri('http://localhost:4000/v1')
#include "pankhllm.h"   /* single .c/.h, libcurl */

char *body = pankh_simple_body("who covers R-3?", NULL, "[\"private\"]", 0);
pankh_response r = pankh_chat("http://localhost:4000", body, NULL);
printf("%d %s\n", r.status, r.body);
pankh_free(r.body); pankh_free(body);
#include "pankhllm.hpp"   // header-only, libcurl

pankhllm::Client pk("http://localhost:4000");
std::string json = pk.chat(R"({"model":"auto","messages":[{"role":"user","content":"who covers R-3?"}]})");
std::cout << json << std::endl;
curl -sS http://localhost:4000/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"auto","messages":[{"role":"user","content":"who covers R-3?"}]}'

# the response carries "pankhllm": {"model": "decision:coverage_owner", "attempts": [...]}
# train on your files, then serve the bundle
python3 notebooks/run_trainer.py --skills ./skills --prompts agent.yaml --logs app.jsonl --teacher ollama
./pankh-work/bundle/run.sh        # http://localhost:4000/v1

# or just the router, from source
cargo install pankhllm && pankhllm serve --port 4000
Demo · a recorded race

Same question. Same executor. Watch who answers first.

Left: the generative planner, a local 12B model, decides and plans. Right: pankhllm's own decision model, trained on 20,000 generated questions, decides in-process. Timings and answers are exactly as recorded on a laptop; the executor returns fictional numbers.

Playback is slowed for the fast lane so you can see it; the numbers are real.

Your LLM · 12B planner, GPU

0 ms
no LLM call

pankhllm · own model, CPU

0 ms
Try your own pankhllm server instead (needs server.cors_origins set for this page's origin)
Measured

Every number comes from a run in the repo.

Laptop, CPU only for pankhllm. Benchmarks, test sets and the steps that got there are in the repository, including where it's weaker.

1743ms
median answer on held-out questions, down from 1,743 ms with the LLM planner. 14 of 14 correct.
0ms
per decision, in-process, on one CPU core
0%
of 2,000 unseen pharma KPI questions answered with no LLM call, 0 wrong
0KB
all decision models, in one portable file
~0s
to train on 20,000 questions, one CPU core, 45 MB RAM
0req/s
server throughput, 17 MB RAM idle, 11 MB binary, no GPU

Where it's weaker: 96.5% precision on phrasing written by a different model. Details in the benchmark and the journal.

The scorecard that matters, from your own traffic

  • p50, p95 and p99 latency before and after: run decisions.mode: shadow and the nightly miner reports both
  • share of requests that bypassed the LLM safely, and the precision of those decisions
  • model calls eliminated and cost saved, per lane
  • behaviour under drift: unfamiliar wording falls back to the LLM by design, failing routes are quarantined
  • how fast coverage grows, night after night

The numbers above are benchmark and recorded runs on a fictional catalog. Production numbers belong to production deployments; shadow mode is how you get yours before acting on anything.

Teach it your domain

Hand over your skills and logs. Get a trained gateway back.

One command, or one sentence to your coding agent. Labels are optional: a teacher LLM labels your logs once, offline.

  • Skills, prompts, logs. Markdown skills, agent YAML, OpenAI request logs, CSV or plain text.
  • Your choice of teacher. Claude, OpenAI, Azure OpenAI / AI Foundry, or a local Ollama model that keeps data on your machine.
  • Ready to run. Server, models.json, config and skills in one bundle.
  • Ask your agent. Ships as a skill for Claude Code and OpenAI Codex: "train pankhllm on these files".

  
How it works

Six lanes. Each one fails open to the next.

pankhllm picks the cheapest path it can prove, per request. Adding a lane can only help: when a lane isn't sure, the next one answers.

LaneWhat answersLatencyLLM calls
Cachean identical question, same tenant and data version~1 ms0
Learned routea question shape seen before, with new entities~1 ms + tool0
Rulea pattern you pinned~1 ms + tool0
Decisionpankhllm's own model picks the operation, tool or skill and fills the parameters0.2 ms + tool0
Plannera new question that maps to your catalog; this is the teacherone small call1
Agentexplanation, judgement, open-ended work: your LLM, hedged and streamedas neededn

What it is

  • A self-learning, OpenAI-compatible gateway in one Rust binary
  • Its own tiny decision models: operations, tools, skills
  • Routing, hedging, caching and budgets around the LLM calls that remain
  • A trace store and overnight miner that turn traffic into training data

What it isn't

  • Not an LLM, and not a replacement for yours
  • Explanations, judgement and open-ended writing still go to your model
  • With only a few dozen logged questions it stays cautious; it takes over as logs grow