Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

go-llm

One clean Go interface to chat LLMs — Anthropic, OpenAI, ChatGPT subscription plans, OpenRouter, and self-hosted servers (vLLM, Ollama, any OpenAI-compatible endpoint) — with normalized streaming, tool calling, structured output, reasoning, usage, cost, and errors.

go-llm is a low-level provider client library: it does one layer well and stops there. No agent loops, no prompt frameworks, no magic. The core package has zero third-party dependencies; official vendor SDKs are used where they exist and pulled in only by the provider package you import.

Highlights

  • One llm.Provider interface — blocking Chat and streaming ChatStream (iter.Seq2 iterators), the same request/response model everywhere, per-provider escape hatches down to the raw SDK client.
  • Providers: Anthropic (API key), OpenAI (Responses API), OpenAI Codex (ChatGPT Plus/Pro subscription OAuth), OpenRouter, and self-hosted vLLM. First-party presets run the shared offline conformance suite, and credentialed presets have capability-driven live scenarios. chatcompletions.New(baseURL) covers any other OpenAI-compatible server (Ollama, llama.cpp, Groq, Together, ...).
  • Subscription auth: OpenAI Codex can mint browser PKCE credentials or consume and auto-refresh compatible existing credentials. llm.LoadAuthFile reads the documented credential union; minted and renewed tokens are returned to your code to persist.
  • Tools: parallel calls, streamed arguments, and a defined contract for malformed tool calls — rescue what's rescuable, drop the rest visibly (ToolCallDropped), with an opt-in llm.RetryDroppedToolCalls middleware. OpenAI-compatible no-argument tool schemas are sent with parameters.properties: {}; non-object properties values fail locally.
  • Structured output: llm.Parse[T] decodes model output straight into your struct (native JSON-schema mode where supported, forced-tool or JSON-mode fallback elsewhere), and schema.For[T] generates JSON Schema from Go types for tool inputs, with schema.ValidateArgs for checking model-emitted arguments.
  • Reasoning: one Effort dial across providers; reasoning output is normalized, and raw provider payloads (signed thinking blocks, encrypted reasoning items) round-trip so same-provider replay just works.
  • Persistence & portability: canonical, versioned JSON for messages and responses. A tool-using conversation started on one provider can be serialized and continued on another — cross-provider handoff is part of the live test matrix, not a hope.
  • Usage & cost: normalized tokens (including cache reads/writes and reasoning tokens), CostUSD from native provider reporting (OpenRouter) or estimated from an embedded, refreshable models.dev price table, including request-wide long-context pricing tiers — with provenance in CostSource (native vs estimated) — plus ContextUsage for context-window accounting. Live model catalogs preserve independent input/output/cache-read/cache-write availability, so unknown rates are not displayed as free while explicit zero rates remain visible.
  • Capabilities: provider-wide Capabilities() discovery with pre-flight request validation, plus advisory per-model capabilities from rich live catalogs such as OpenRouter and Codex. Unsupported provider features fail fast with ErrUnsupported; incomplete model metadata never blocks a call.
  • Conveniences: Session (auto-managed history, session tools with AddToolResults + Continue, cumulative usage), History, PromptTemplate, llm.Ptr for optional scalar fields, middleware via llm.Wrap, and observability built in but silent by default (slog logging, UsageTracker, wire capture via WithWireCapture). Sensitive headers are always redacted; captured URLs and bodies remain application-sensitive data. Built-in failure logs use safe provider-error summaries and never emit provider-controlled bodies or metadata.
  • Testing: llmtest — like net/http/httptest, but for code that consumes go-llm.
  • CLI: llm-cli, a curl-like frontend built entirely on the public API.

Install

go get github.com/pkieltyka/go-llm

Go 1.26 is the minimum version for users of the module. Development and releases use Go 1.27.0 or newer, while CI retains a Go 1.26 compatibility lane. Provider SDK dependencies are pulled only when importing provider packages.

Quick start

package main

import (
	"context"
	"fmt"

	llm "github.com/pkieltyka/go-llm"
	"github.com/pkieltyka/go-llm/providers/openai"
)

func main() {
	ctx := context.Background()

	p, err := openai.New() // reads OPENAI_API_KEY
	if err != nil {
		panic(err)
	}

	resp, err := p.Chat(ctx, &llm.Request{
		Model:    "gpt-5.5",
		Messages: []llm.Message{llm.UserText("Explain Go iterators in one sentence.")},
	})
	if err != nil {
		panic(err)
	}

	fmt.Println(resp.Text())
}

Streaming:

for text, err := range llm.StreamText(p.ChatStream(ctx, req)) {
	if err != nil {
		return err
	}
	fmt.Print(text)
}

Structured output:

type Summary struct {
	Title string   `json:"title"`
	Tags  []string `json:"tags"`
}

summary, resp, err := llm.Parse[Summary](ctx, p, &llm.Request{
	Model:    "gpt-5.5",
	Messages: []llm.Message{llm.UserText("Summarize this release note.")},
})
_, _ = summary, resp

Providers

Package Auth Notes
providers/anthropic ANTHROPIC_API_KEY or WithAPIKey Messages API
providers/openai OPENAI_API_KEY or WithAPIKey Responses API — reasoning survives across tool-call turns
providers/openaicodex WithOAuth only (ChatGPT Plus/Pro) Responses wire shape; explicit Models calls use cached authenticated discovery with a curated fallback
providers/openrouter OPENROUTER_API_KEY or WithAPIKey Chat Completions; routing/plugins, explicit rich model discovery, and native per-request cost reporting
providers/vllm optional WithAPIKey (vLLM --api-key) Self-hosted vLLM preset: host-first, era-aware, live-tested
providers/ollama none Data-only local-Ollama preset over the engine below (community-verified)
providers/chatcompletions optional WithAPIKey Public engine for ANY OpenAI-compatible server: New(baseURL, ...) + declarative Compat quirks
anthropic.New()  // ANTHROPIC_API_KEY
openai.New()     // OPENAI_API_KEY
openrouter.New() // OPENROUTER_API_KEY

OpenAI Responses callers can select the user-visible reasoning-summary detail without adding an OpenAI-specific field to llm.Request:

resp, err := p.Chat(ctx, &llm.Request{
	Model:    "gpt-5.5",
	Effort:   llm.EffortHigh,
	Messages: []llm.Message{llm.UserText("Compare the two proposals.")},
	ProviderOptions: openai.Options{
		ReasoningSummary: openai.ReasoningSummaryConcise,
	},
})

The supported selectors are ReasoningSummaryAuto, ReasoningSummaryConcise, and ReasoningSummaryDetailed. Leaving the field empty preserves the existing behavior (summary: "auto" when Effort is set); unknown values fail before a request is sent. OpenAI remains authoritative about which models accept each selector. This option is not available on the separate openaicodex subscription provider.

OpenAI Codex consumes credentials minted by compatible existing tools, refreshes them automatically, and hands renewals back to your code to persist; the root library never writes credential files itself:

auth, err := llm.LoadAuthFile("auth.json") // documented credential union
if err != nil {
	panic(err)
}

codex, err := openaicodex.New(openaicodex.WithOAuth(auth["openai-codex"], func(ctx context.Context, updated llm.AuthCredential) error {
	return persistCredential(ctx, updated)
}))

Codex can also mint the initial credential through its provider-owned browser PKCE flow. Begin returns a redacting loopback launch target; Complete waits for the automatic callback, while Submit accepts a copied complete callback URL/query if the browser cannot reach localhost. The host must store the returned credential before passing it to WithOAuth:

flow, err := openaicodex.NewLoginFlow()
authorization, err := flow.Begin(ctx)
showLoginURL(authorization.URL())
credential, err := flow.Complete(ctx)
err = persistCredential(ctx, credential)

The callback binds only 127.0.0.1, uses the registered localhost path on port 1455 with port 1457 as its fallback, and never persists the credential. Live OAuth verification is still required because automated tests cannot complete an account login. Device authorization and a built-in CLI login UX remain deferred.

Persistence callbacks must honor their context and return only after the rotated credential is durably stored; an error prevents the provider from publishing it. A credential containing a refresh token requires a non-nil callback; access-only credentials may pass nil. To deliberately keep rotations only in memory, pass an explicit context-aware no-op:

discardRotation := func(ctx context.Context, _ llm.AuthCredential) error {
	if err := ctx.Err(); err != nil {
		return err
	}
	return nil
}

This can leave the stored refresh token stale after restart and should be a conscious application decision. Provider-specific request extensions live in each provider's Options type, passed through Request.ProviderOptions; the raw SDK client is always reachable via each provider's Client(). Ordinary openai.Options fields use go-llm and standard-library types; only advanced escape hatches such as Provider.Client, chatcompletions.Dialect, chatcompletions.Config, and Provider.BuildParams are vendor-coupled and stability-exempt before v1.

Self-hosted (vLLM, Ollama, any OpenAI-compatible server)

Self-hosted servers are first-class: constructors are host-first and the API key is optional (no environment fallback). The providers/vllm preset targets the current stable vLLM v0.26.0 protocol and knows its dialect — reasoning output (parsed to llm.ReasoningPart, streamed as ReasoningDelta), Effort → reasoning_effort, choice-less mid-stream error events, max_model_len as ModelInfo.ContextWindow, and typed extensions (top_k, min_p, stop_token_ids, chat_template_kwargs, thinking_token_budget, vllm_xargs, plus native structured_outputs constraint modes — regex/choice/grammar/structural-tag via vllm.StructuredOutputs; JSON schema stays on the unified ResponseFormat):

p, err := vllm.New("http://localhost:8000/v1") // keyless by default
if err != nil {
	panic(err)
}

model, err := p.ResolveModel(ctx, "qwen") // resolves deployment prefixes/suffixes
if err != nil {
	panic(err)
}

resp, err := p.Chat(ctx, &llm.Request{
	Model:    model.ID,
	Effort:   llm.EffortNone, // thinking-by-default models answer tersely
	Messages: []llm.Message{llm.UserText("hello")},
	ProviderOptions: vllm.Options{
		XArgs: map[string]any{"custom_engine_arg": "1"},
	},
})

For deployments that support it, vllm.Options.ThinkingTokenBudget caps reasoning without adding a provider-neutral knob. When MaxTokens is set, go-llm clamps the configured budget to reserve 1,024 tokens for the visible answer and rejects thinking-disabled combinations before sending the request.

vLLM also gives you exact token counting: the preset exposes the server's /tokenize endpoints as typed extension methods, so context accounting is ground truth (server-rendered chat template, tools included) instead of an estimate — live-verified to match a real request's prompt_tokens exactly:

result, err := p.Tokenize(ctx, req)     // same conversion + validation as Chat
usage := result.ContextUsage()          // exact count vs the model's max_model_len
fmt.Printf("prompt occupies %d of %d tokens (%.2f%%)\n",
	usage.UsedTokens, usage.Window, usage.UsedPercent)

Detokenize and TokenizerInfo round out the family. For everything else that speaks the Chat Completions shape, the same engine is public:

// Ollama's local convention, as a data-only preset:
p, err := ollama.New("") // http://localhost:11434/v1

// Or any OpenAI-compatible endpoint, with quirks declared as data:
p, err = chatcompletions.New("https://api.example.com/v1",
	chatcompletions.WithName("example"),
	chatcompletions.WithAPIKey(os.Getenv("EXAMPLE_API_KEY")),
	chatcompletions.WithCompat(chatcompletions.Compat{StreamIncludeUsage: true}),
)

Bonus recipe: current vLLM v0.26.0 also serves an Anthropic /v1/messages endpoint, so the Anthropic provider can target a vLLM box directly — anthropic.New(anthropic.WithBaseURL("http://localhost:8000"), anthropic.WithAPIKey("dummy")) completes chats with usage and even maps Qwen thinking output to reasoning parts (vLLM emits Anthropic-style thinking blocks there).

Persistence and cross-provider handoff

llm.MarshalMessages / llm.UnmarshalMessages give you a canonical, versioned JSON envelope for conversation history — safe to store in a database and reload across releases. Because every adapter re-encodes the neutral history into its own wire shape (and drops what only the original provider can accept, like signed thinking blocks), a saved conversation can be continued on a different provider: mid-conversation failover, cheap-model tool loops handed to a stronger model for synthesis, or histories that simply outlive your vendor choice.

CLI

go install github.com/pkieltyka/go-llm/cmd/llm-cli@latest
llm-cli -p openai -m gpt-5.5 "write a short haiku about Go"
echo "long input" | llm-cli -p anthropic -m claude-opus-4-8 -s "summarize stdin"
llm-cli -p openrouter -m openai/gpt-5.5 --usage --json "return a JSON status"
llm-cli -p openai-codex --auth-file ~/.config/llm/auth.json -m gpt-5.4 "review this diff"

llm-cli models -p openrouter          # list models (table or --json)

llm-cli -p openai -m gpt-5.5 --save chat.json "Start a checklist"
llm-cli -p anthropic -m claude-opus-4-8 --load chat.json --save chat.json "Continue it"

llm-cli models explicitly performs provider model-discovery network I/O. OpenRouter uses exactly its /models response and does not fall back to models.dev. Listings include advisory supported/default reasoning effort, whether the provider positively reports reasoning as required, and model capabilities. Text output shows EFFORTS, DEFAULT EFFORT, and REASONING REQUIRED; --json emits the corresponding optional fields. ReasoningRequired is selection metadata only and never overrides the caller's Request.Effort.

The last pair saves a conversation with one provider and continues it with another — the handoff described above, from the shell.

For OpenAI Codex, authentication precedence is --auth-file, then OPENAI_CODEX_ACCESS_TOKEN, then compatibility --api-key. Prefer the first two forms: command-line API-key values are visible in process arguments and often shell history. Auth files are loaded only when explicitly requested. Refresh publication holds a bounded cross-process advisory lock across the full read/modify/atomic-write transaction so concurrent CLI processes do not lose one another's credential updates.

Testing your code

Use llmtest for unit tests that should never contact a provider — like net/http/httptest, but for code that consumes go-llm:

p := llmtest.New(llmtest.WithCapabilities(llm.CapabilityJSONSchema))
p.EnqueueResponse(&llm.Response{Parts: []llm.Part{llm.Text(`{"ok":true}`)}})

Script responses, streams, and errors; assert on the requests your code made via p.Requests(). See examples/testing for a complete worked example.

Provider authors should run two complementary offline suites. RunConformance checks the common lifecycle, streaming, and normalized-result contract. RunCapabilityConformance checks that each reviewed advertised capability activates its exact native request fields and returns the expected normalized result. Its CapabilityInvocation is fixture control data passed to an isolated provider factory; it never changes the real request or its model. Explicit profile exemptions mean “offline evidence gap,” not “unsupported.” Neither suite proves live account/model availability, quota, cache admission, or service reliability; credentialed e2e scenarios remain the live evidence. The advanced chatcompletions.NewWithDialect conformance suite intentionally stays base-only; the public generic engine owns the reusable compatible-engine activation profile.

Repository commands:

make build       # compile all packages and create bin/llm-cli
make test        # offline tests with the race detector
make check       # complete credential-free, network-dependent CI gate
make models      # refresh the embedded root models.json snapshot and provenance

Live end-to-end tests are behind the live build tag and read credentials from gollm-test.json (copy gollm-test.json.sample; missing credentials skip visibly, never fail):

make e2e-test

Examples

Every program in examples/ is dual-mode: it runs offline out of the box against the scripted llmtest provider, and switches to the real API when ANTHROPIC_API_KEY or OPENROUTER_API_KEY is set.

go run ./examples/chat                        # offline, scripted fallback
ANTHROPIC_API_KEY=... go run ./examples/chat  # same program, real API
Example Shows
chat One blocking request/response
stream Streaming deltas with llm.StreamText
tools A full tool round trip: call → local execution → result → answer
parse llm.Parse[T] structured output into a Go struct
history-replay Serialize a conversation and continue it on another provider (a real Anthropic→OpenRouter handoff when both keys are set)
observability UsageTracker, logging, and cost via llm.Wrap middleware
provider-selection Choosing a provider at runtime: env keys, an LLM_AUTH_FILE credential file (including codex OAuth), or the offline fake
testing Unit-testing your code with llmtest: request assertions, streaming, error paths (go test ./examples/testing/)

Godoc examples in example_test.go run offline as part of the test suite.

Status

Pre-1.0: public APIs are intended to be small and stable, but breaking changes may happen before v1.0.0 when provider behavior or the unified model needs correction (the chatcompletions.Dialect interface is explicitly an advanced, stability-exempt surface — prefer the declarative Compat). After v1.0.0, standard Go module compatibility rules apply.

License

MIT.

About

Unified LLM provider library in Go across Anthropic, OpenAI, vLLM, OpenRouter and many more

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages