pankhllm started as an internal fix for LLM latency in production. Our agents kept paying a large model for the same kinds of decisions: which tool, which skill, which report. pankhllm learns those recurring decisions from your traffic, makes the safe ones itself in 0.2 ms on a CPU, and leaves everything unfamiliar to the LLM you already use.
Your LLM keeps answering while pankhllm watches. Every decision it sees becomes a label; every night it retrains; every morning more questions never reach the LLM.
Planner choices, clean tool calls and explicit skill picks are recorded. Only the question's shape is kept.
pankhllm mine retrains the models and picks a confidence threshold for your target precision.
It acts only when confident, the wording is familiar, and the question fills the operation's parameters exactly.
Nothing breaks while it learns. That answer becomes the next label.
LangChain, Semantic Kernel, Microsoft Agent Framework, the Vercel AI SDK, or any OpenAI client. Small pankhllm clients ship for Python, TypeScript, JavaScript, C#, Go, Java, Rust, Ruby, PHP, C, C++ and the shell.
from openai import OpenAI client = OpenAI(base_url="http://localhost:4000/v1", api_key="...") # was: your LLM provider r = client.chat.completions.create(model="auto", messages=[{"role": "user", "content": "who covers R-3?"}]) # r.model_extra["pankhllm"]["model"] == "decision:coverage_owner" (2 ms, no LLM call) # or the pankhllm client: pip install pankhllm from pankhllm import Client a = Client("http://localhost:4000").ask("who covers R-3?", tags=["private"]) print(a.text, a.model)
import OpenAI from "openai"; const client = new OpenAI({ baseURL: "http://localhost:4000/v1", apiKey: "..." }); const r = await client.chat.completions.create({ model: "auto", messages: [{ role: "user", content: q }] }); // or the pankhllm client: npm install @pankhllm/client import { PankhClient } from "@pankhllm/client"; const a = await new PankhClient("http://localhost:4000").ask(q, { tags: ["private"] });
// Node 18+, plain JavaScript, no package needed const r = await fetch("http://localhost:4000/v1/chat/completions", { method: "POST", headers: { "content-type": "application/json" }, body: JSON.stringify({ model: "auto", messages: [{ role: "user", content: "who covers R-3?" }] }) }); const v = await r.json(); console.log(v.choices[0].message.content, v.pankhllm.model); // "decision:coverage_owner" // with the official openai package (CommonJS) const OpenAI = require("openai"); const client = new OpenAI({ baseURL: "http://localhost:4000/v1", apiKey: "..." }); // or the pankhllm client: npm install @pankhllm/client import { PankhClient } from "@pankhllm/client"; const a = await new PankhClient("http://localhost:4000").ask("who covers R-3?");
// Semantic Kernel / Microsoft Agent Framework: change the endpoint, keep your plugins and loop var kernel = Kernel.CreateBuilder() .AddOpenAIChatCompletion("auto", new Uri("http://localhost:4000/v1"), apiKey: "...") .Build(); // or the pankhllm client: dotnet add package Pankhllm var pk = new Pankhllm.PankhClient("http://localhost:4000"); var a = await pk.AskAsync("who covers R-3?", new AskOptions { Tags = ["private"] }); Console.WriteLine($"{a.Text} via {a.Model}");
import "github.com/singhpratech/pankhllm/sdks/go" pk := pankhllm.New("http://localhost:4000") a, err := pk.Ask(ctx, "who covers R-3?", pankhllm.Options{Tags: []string{"private"}}) if err != nil { log.Fatal(err) } fmt.Println(a.Text, a.Model) // a.Model == "decision:coverage_owner" // any OpenAI Go client works too: set the base URL to http://localhost:4000/v1
import dev.pankhllm.PankhClient; var pk = new PankhClient("http://localhost:4000"); var a = pk.ask("who covers R-3?", new PankhClient.Options().tags(List.of("private"))); System.out.println(a.text() + " via " + a.model()); // Spring AI / LangChain4j / the official OpenAI Java SDK: base URL http://localhost:4000/v1
use pankhllm_client::{Client, Options}; let pk = Client::new("http://localhost:4000"); let a = pk.ask("who covers R-3?", &Options { tags: vec!["private".into()], ..Default::default() }).await?; println!("{} via {}", a.text, a.model); // the router itself: cargo install pankhllm && pankhllm serve --port 4000
require "pankhllm" pk = Pankhllm::Client.new("http://localhost:4000") a = pk.ask("who covers R-3?", tags: ["private"]) puts "#{a.text} via #{a.model}" # ruby-openai works too: OpenAI::Client.new(uri_base: "http://localhost:4000/v1")
use Pankhllm\Client; $pk = new Client('http://localhost:4000'); $a = $pk->ask('who covers R-3?', ['tags' => ['private']]); echo $a['text'], ' via ', $a['model']; // openai-php works too: OpenAI::factory()->withBaseUri('http://localhost:4000/v1')
#include "pankhllm.h" /* single .c/.h, libcurl */ char *body = pankh_simple_body("who covers R-3?", NULL, "[\"private\"]", 0); pankh_response r = pankh_chat("http://localhost:4000", body, NULL); printf("%d %s\n", r.status, r.body); pankh_free(r.body); pankh_free(body);
#include "pankhllm.hpp" // header-only, libcurl pankhllm::Client pk("http://localhost:4000"); std::string json = pk.chat(R"({"model":"auto","messages":[{"role":"user","content":"who covers R-3?"}]})"); std::cout << json << std::endl;
curl -sS http://localhost:4000/v1/chat/completions \ -H 'content-type: application/json' \ -d '{"model":"auto","messages":[{"role":"user","content":"who covers R-3?"}]}' # the response carries "pankhllm": {"model": "decision:coverage_owner", "attempts": [...]}
# train on your files, then serve the bundle python3 notebooks/run_trainer.py --skills ./skills --prompts agent.yaml --logs app.jsonl --teacher ollama ./pankh-work/bundle/run.sh # http://localhost:4000/v1 # or just the router, from source cargo install pankhllm && pankhllm serve --port 4000
Left: the generative planner, a local 12B model, decides and plans. Right: pankhllm's own decision model, trained on 20,000 generated questions, decides in-process. Timings and answers are exactly as recorded on a laptop; the executor returns fictional numbers.
Playback is slowed for the fast lane so you can see it; the numbers are real.
server.cors_origins set for this page's origin)Laptop, CPU only for pankhllm. Benchmarks, test sets and the steps that got there are in the repository, including where it's weaker.
Where it's weaker: 96.5% precision on phrasing written by a different model. Details in the benchmark and the journal.
decisions.mode: shadow and the nightly miner reports bothThe numbers above are benchmark and recorded runs on a fictional catalog. Production numbers belong to production deployments; shadow mode is how you get yours before acting on anything.
One command, or one sentence to your coding agent. Labels are optional: a teacher LLM labels your logs once, offline.
models.json, config and skills in one bundle.pankhllm picks the cheapest path it can prove, per request. Adding a lane can only help: when a lane isn't sure, the next one answers.