Changelog
Release notes pulled live from the project changelog - the notable changes per version, shorter than the raw commit history.
v0.6.1
LatestAdded- Added GPT-6 Astra support across tool search, additional tools, long-context pricing, xhigh and max thinking levels, and its explicit thinking-level map.
- Added five-times-faster mouse wheel scrolling while holding Alt in fullscreen mode.
Changed- Changed the generated image model catalog to the current OpenRouter listing.
- Changed the built-in read, write, edit, and bash tools to request strict JSON-schema sampling by default instead of only under
KNIGHTCODE_EXPERIMENTAL. - Changed fullscreen scrollbars to render muted thin tracks with contrasting proportional two-cell-minimum thumbs, reserve an unstyled column in
alwaysmode, reveal hiddenautotracks on pointer entry, expand the same-colored thumb on hover, and support track-click jumping in addition to thumb dragging, with optionalscrollbarTrackandscrollbarThumbtheme colors falling back to muted and text. - Changed fullscreen transcript search to cache unchanged results, index ASCII runs, and highlight only visible matches, so latency no longer grows with transcript size.
- Changed clipboard handling to use small built-in macOS, Windows, and X11 native helpers instead of an external dependency, running native reads on worker threads and making the command-line fallbacks (
pbcopy,clip.exe,wl-copy,xclip) asynchronous. Incremental X11 transfers, legacy text encodings, and native image formats are preserved.
Removed- Removed Grok Build 0.1 from the built-in xAI model catalog.
Fixed- Fixed processes killed by a signal reporting success; they now map to a 128 + signal exit code.
- Fixed compiled binaries shipping without the TUI's native helpers, which left clipboard reads on the command-line fallbacks and dropped Shift+Tab on Windows. Each target's prebuilds are now copied next to the executable.
- Fixed
fdfailing to start on musl-based Linux distributions by downloading the statically linked musl builds of bothfdandripgrep. - Fixed post-login model selection for Radius, whose per-account catalog is empty until the first authenticated refresh; selection now waits for that refresh, defaults to
balanced, and falls back to catalog order. - Fixed the model, scoped-model, and thinking selectors hardcoding Ctrl+S to save; the shortcut is now the
app.models.saveandapp.thinking.savekeybindings and the on-screen hint follows a rebind. - Fixed mouse hover changing selection and recentering autocomplete and settings lists, causing clicks to target a different item.
v0.6.0
Added- Added per-turn thinking effort preservation for Claude models, so a conversation whose turns were answered at different effort levels no longer replays as if every turn used the current one.
Changed- Changed the agent engine to a lane-owned durable architecture, rebuilding the harness, session, protocol, client and server layers in one pass.
- Added
@knightcode/chord, the application composition runtime (services, replicated state, RPC, plugins) the harness, protocol, client and server now build on. - Changed the harness runtime to a lane-owned drive: durable execution primitives, an effect gate, hooks, restore, and a mutation line replace the earlier operation-task/procedure runtime, and the transitional
runtime2andrestorelayers are gone. - Changed the session layer to bound values and lists with a commit/fork/mutation-line model, storage and repository conformance suites, and session benchmarks; the SQLite backend follows it.
- Changed the protocol, client and server to Chord-routed services: a single
protocol.ts, a hosted harness manager, server identities, session directories, draining, and the wrong-server/session-not-found error set replace the earlier RPC schemas and live-session manager. - Changed the experimental CLI to durable server and client commands, with
--provider,--model,-e,--continue,--resume,--server-idand--session-dir, dropping the local demo runtime and session-worker process. - Changed built-in tool rendering to load from
core/tools/renderers/, so a process that only displays tool output no longer pulls in the execution path; the KnightCode tool-output style (Read(...),Search(...),Update(...), collapsed one-line summaries) moved with it. - Added click-to-expand on tool results, wired through the gutter shell so the bullet and continuation-marker layout keeps working.
- Fixed the reserved fork namespace guard, which tested for a stale prefix and so never matched a
knightcode.namespace. - Fixed Windows portability across the session, socket and CLI-spawning suites:
node --importnow receives a file URL, session tests resolve their workspace paths, and the Unix-socket suites are skipped where the platform cannot bind them.
Fixed- Fixed
knightcode configignoring theshowHardwareCursorandclearOnShrinkterminal settings.
v0.5.4
Added- Added an entries argument to
SessionManager.inMemory(), so an SDK embedder can resume a session held outside the filesystem — in a database, say — without writing it to a temporary.jsonlfile first. - Added the relational algebra join operators to LaTeX rendering:
\bowtie,\Join,\ltimes,\rtimes,\leftouterjoin,\rightouterjoinand\fullouterjoin. - Added a
vllmPrioritycompat flag for customopenai-completionsproviders. Set it on a model and requests carry a top-levelpriorityfield, which a vLLM server running with--scheduling-policy priorityuses to order work; lower values are served first. Unset by default, so nothing changes for providers that do not want it.
Changed- Changed the branch summary output cap from 2048 to 4096 tokens, clamped to the model's own limit, so summaries of long branches are no longer cut off mid-sentence.
- Changed the Cloudflare AI Gateway binding transport to pass requests straight to the Workers AI binding's
fetchrather than translating them into universal-endpoint calls.createGatewayBindingFetchis replaced bycreateAiBindingFetch(env.AI), which supports every method, non-JSON bodies and streaming request bodies instead of rejecting them. - Changed the bundled model catalog to a fresh regeneration from models.dev. GitHub Copilot drops eight models that the provider no longer serves (
claude-opus-4.5,claude-opus-4.6,claude-sonnet-4,claude-sonnet-4.5,gemini-3.1-pro-preview,gpt-4.1,gpt-5.2,gpt-5.2-codex) and gainsclaude-fable-5.1andgemini-3.8-flash. This regeneration is also what activates the Copilot Fable 5 routing fix, which changed only the generator and so never reached the committed data. Baseten gainszai-org/GLM-5.3-Fast, Cloudflare AI Gateway gainsclaude-fable-5.1, and OpenCode Go gainsomen-alpha. - Removals only take effect through regenerated data: the remote catalog overlay merges by id and can add or update models, but never removes them, so a model that disappears upstream keeps appearing until the bundled catalog is refreshed.
- Changed the selectors in
/thinking,/model,/scoped-models,/trustand per-model thinking settings to keep the active option marked while browsing, by moving the marker into a fixed column ahead of the label./scoped-modelsnow uses the same per-item toggle as the rest, strikes through models that are no longer available, and no longer collapses to a single model when the first one is toggled off. - Changed the theme settings selectors to keep the configured theme marked while browsing, matching the other selectors. Both the fixed-theme list and the light/dark lists behind Automatic now show the marker in a fixed column.
- Changed the streaming working indicator to render in the editor's top border instead of on its own row above it, so the editor no longer shifts up and down as a turn starts and finishes. It picks up the editor's border colour, which already tracks the thinking level. Custom editors from extensions keep the standalone row unless they opt in with
embedWorkingStatus.
Fixed- Fixed aborting a session leaving an in-progress compaction or branch summary running. Escape during
/compact, or an RPCabort, now cancels it and waits for the session to actually be idle before returning. - Fixed Baseten's GLM-5.2 and GLM-5.2-Fast being advertised as accepting images. The catalog reports image input for them but the endpoints are text-only, so attaching an image produced a provider error instead of being caught up front.
- Fixed a Codex response being dropped when the server closed the stream without a blank line after the final event. The last frame is now processed at EOF instead of being discarded with the buffer.
- Fixed GitHub Copilot Claude Fable models being served through the OpenAI completions adapter, which dropped the selected reasoning level. They now route through the Anthropic Messages adapter like the other Claude 4.x and 5.x models on that provider.
- Fixed Fireworks GLM models other than GLM-5.2 being served through the Anthropic-compatible endpoint, which does not accept them. Every
glm-model on Fireworks now uses the OpenAI completions endpoint, so GLM-5.3 and GLM-5.3 Flash work. - Fixed forking a compacted session losing the messages after the compaction boundary when that boundary pointed at a label. Labels are dropped from the forked path, which left the boundary pointing at an entry that no longer existed.
- Fixed importing a session file silently overwriting a stored session that happened to have the same filename. The import is now written alongside it under a numbered name.
- Fixed a proxied request hanging when the server closed the stream without sending a terminal event, and a final event that arrived without a trailing newline being dropped. The first now surfaces as an error, the second is processed.
- Fixed proxied plain-HTTP provider requests hanging after a tool call by tunneling them with CONNECT again, restoring the behaviour Undici changed in 8.7.
- Fixed Qwen3.8 Flash offering the wrong thinking levels on the Qwen Token Plan providers: it advertised high and max, which it does not accept, instead of low, medium and xhigh. It is also now listed on the Individual plan, where it is available.
- Fixed the built-in tools ignoring the working directory supplied on the extension context.
read,write,edit,ls,find,grepand the shell tool resolved relative paths against the directory captured when the tool was created, so a caller running a tool against a different directory operated in the session's directory instead of its own. - Fixed
fdandripgrepfailing to download behind shared egress IPs, where the anonymous GitHub API rate limit is permanently exhausted. The latest release is now resolved from the release page redirect, which costs no API quota. A failed download also reports the underlying network error instead of a bare "fetch failed". - Fixed
knightcode updatereporting every install as a standalone binary.bin/knightcodespawns the compiled binary out ofnode_modules, so install detection now classifies a binary by where it sits rather than by how it was built, and moves the running executable aside on Windows so npm can replace it. - Fixed the write tool reporting UTF-16 code-unit counts as byte counts by removing the misleading count from its result.
- Added an entries argument to
v0.5.3
Fixed- Run the auto-compaction threshold check between turns of an agent run, so a tool batch that fills the context window is compacted before the next assistant request instead of overflowing it.
- Add a
fullscreenCopyOnSelectsetting (defaulttrue). Turn it off and a fullscreen mouse selection stays highlighted instead of being copied on mouse release, andCtrl+Xcopies the active selection rather than the last assistant message. - Settle the running turn before an in-memory
/fork, so the aborted assistant message and its tool results are no longer appended to the freshly forked session. - Merge Mistral streaming tool-call chunks by their
index, so a call whose id and name arrive only in the first chunk is no longer split into two tool calls with truncated arguments. A name that arrives on a later chunk is picked up rather than left empty. - Match
NO_PROXYentries against the root domain and its subdomains, and parse IPv6 hosts andhost:portentries correctly, so a bareexample.comentry also bypasses the proxy forapi.example.comandnotexample.comno longer matches it. A bare*entry now bypasses everything even when listed alongside other entries, and an entry with a malformed port is dropped rather than widened into a host-wide bypass. - Add a
supportsMaxOutputTokenscompat flag foropenai-responsesmodels (defaulttrue). Set it tofalsefor a gateway that rejectsmax_output_tokensand the parameter is omitted instead of failing the request. - Stop already-prepared tool calls from running when a parallel batch is aborted during preflight, so cancelling at a permission prompt no longer lets the remaining tools in that batch execute.
- Ignore a failing SIGWINCH self-signal at terminal startup, so sandboxes whose seccomp or LSM policy denies
kill(2)no longer crash on launch. The dimension refresh is skipped instead. - Tidy the tool call transcript block. The dark theme's
greenandrednow hold the pinned diff hexes, so the success bullet,✓marks, bash mode and markdown code blocks match the diff colours instead of staying olive. Shell tool call headers are clamped to a single line — a long command no longer wraps several rows of quoted URL over the transcript — and the bash expand hint follows its output rather than preceding it, matching every other tool renderer. Line counts in the expand hints are pluralised. - Detect Zed's integrated terminal so it gets truecolor and hyperlinks instead of falling through to the conservative default, and document the Zed key bindings needed for
Shift+Enterand friends.
v0.5.2
Fixed- Read EXIF orientation from JPEGs whose first APP1 segment holds XMP instead of EXIF. Such images previously rendered unrotated.
- Clear a delivered steering or follow-up message that carried only images. The entry previously stayed in the queue forever, leaving the pending count wrong and the message re-queued.
- Give each
/shareits own temp directory so two shares running at once no longer overwrite each other's export or delete the other's file mid-upload. - Keep skills in the system prompt when
readis disabled but a shell tool is available, and tell the model to loadSKILL.mdwithbash(or PowerShell) instead. Skills previously vanished entirely from bash-only tool setups. - Ignore Kitty image conversions that land after the tool image at that position changed, so a streamed partial image no longer replaces the final result.
- Add AgentRouter as a built-in provider.
AGENTROUTER_API_KEYenables five AgentRouter models, defaulting toagentrouter/glm-5.3, with Claude routed through the Anthropic Messages endpoint and the rest through the OpenAI-compatible one. Token prices come from AgentRouter's rate table rather than upstream list prices; cache costs remain estimates because AgentRouter does not publish its cache ratios.
v0.5.1
Fixed- Tool calls now render as blocks in the transcript. Each one shows a
Bash(...)/Read(...)/Update(...)header with its result collapsed underneath on a⎿gutter, instead of the flat before/after dump. The rest of the chrome — boxed messages, the rounded input frame, the braille spinner, the banner and footer — is unchanged. Fixed a context-window overflow loop. Messages with no provider usage yet are estimated at 4 chars/token, but real tokenizers land nearer 3 on code and JSON, so reservingmax_tokensagainst the raw estimate could push prompt +max_tokenspast the window. The provider rejected it as an overflow, the agent compacted,max_tokensre-expanded into the freed room, and the next request failed the same way. The estimated part is now padded so the reservation stays inside the window.
- Tool calls now render as blocks in the transcript. Each one shows a
v0.5.0
A rebuilt agent core. The agent loop, session storage, provider layer and terminal UI were all replaced. What that buys: Distribution is unchanged — a self-contained compiled binary per platform, no Bun or Node needed at runtime.
Added- A measured ~1,100-token floor for the system prompt and tool definitions — every request is smaller, on every model.
- Real multi-provider support: Anthropic, OpenAI/Codex, OpenRouter, Amazon Bedrock, xAI, Kimi, GitHub Copilot, and any custom endpoint through
models.json. OAuth sign-in where the provider supports it, API keys everywhere else. - Sessions you can leave and come back to: resume, fork, branch, search, and automatic compaction when a conversation outgrows the context window.
- Extensions, skills and prompt templates, discovered from the project or installed globally.
- Headless mode:
--printwithtext,jsonorrpcoutput, for scripting and for driving KnightCode from another program.
v0.4.1
Fixed- Re-inject the current todo list after each tool round so the model's plan stays in context during long turns. Only fires when the list has unfinished items and has changed since the last round.
v0.4.0
Harness reliability: safer edits and recovery from flaky model streams.
Added- No blind or stale edits. A file must be read before it can be edited, and an edit is rejected if the file changed on disk since that read — so a write can't silently clobber newer changes. The read state is rebuilt from the transcript, so it survives a session resume.
- No accidental repeats. Identical read-only tool calls in one round run once instead of duplicating, and the loop guard stops repeated identical calls sooner.
- Auto-retry on flaky streams. Transient stream failures and empty responses retry with exponential backoff (honoring
Retry-After); cancelling mid-backoff no longer fires an extra model call. - Tool errors self-correct. An invalid tool call no longer ends the turn — the model gets the error back and can fix it.
- No misleading diffs. An edit diff shows only after the edit actually applies; failed or rejected edits don't render one.
v0.3.1
Fixed- Fix the
/exitcommand freezing the terminal in packaged builds. Process cleanup usedspawnSync(process.execPath, ["-e", ...])as a sleep, but in a compiled standalone binaryprocess.execPathis the CLI itself, so it relaunched the TUI and blocked forever. Replaced it with an in-process sleep and made exit terminate the process explicitly.
- Fix the
v0.3.0
Add automatic skill discovery, hot-reload, and path-scoped skills so installed skills surface and get loaded without having to be named explicitly.
Added- Skill auto-discovery. Each turn a cheap side-query compares your request against the installed skills, surfaces the relevant ones, and directs the model to load them via the
Skilltool before responding. Surfaced skills appear as a visible↳ Relevant skills: …line in the chat. Controlled by theskills.autoDiscoversetting (on by default). - Skill hot-reload. A file watcher picks up added, edited, or removed
SKILL.mdfiles mid-session, so changes take effect without restarting. Controlled by theskills.hotReloadsetting (on by default). - Path-scoped (conditional) skills. A skill with a
pathsfrontmatter glob is kept out of the always-on skill list and surfaces only when you edit a file matching its globs.
Changed- The skill index injected into the system prompt is now size-bounded: descriptions are truncated to fit the budget and, in the extreme, the listing falls back to names only — but every skill name is always shown, so no installed skill becomes undiscoverable.
- Skill auto-discovery. Each turn a cheap side-query compares your request against the installed skills, surfaces the relevant ones, and directs the model to load them via the
v0.2.1
Added- Memory follow-ups: feed recent tool usage into the recall selector as an extra relevance signal, frame extraction's "new messages" window from a per-session cursor (so durable facts mentioned during gate-skipped turns are still reconsidered), and drain any in-flight memory extraction on
/exit(bounded) so a save isn't dropped at shutdown. - Refresh the supported model catalog with new OpenRouter models:
nvidia/nemotron-3-ultra-550b-a55b:free(Nemotron 3 Ultra 550B),nex-agi/nex-n2-pro:free(Nex N2 Pro),qwen/qwen3.7-plus(Qwen3.7 Plus),z-ai/glm-5.2(GLM 5.2), andmoonshotai/kimi-k2.7-code(Kimi K2.7 Code). Newqwenandnexmodel aliases accompany them. - Accurate per-session cost: enable OpenRouter usage accounting (
usage.include) so each request returns its actual cost. The in-app/cost"Session cost" now sums real costs (correct for free/cached/uncurated models) and only falls back to the local price table when a message has no reported cost. - Session grouping on OpenRouter: send the session id as the
x-session-idheader so a session's requests are grouped in OpenRouter's logs (Sessions tab) and routed stickily to the same provider for better prompt-cache hits. Requests are also tagged with the session id via theuserfield for per-request "Client User ID" attribution.
Changed- Default model is now
nvidia/nemotron-3-ultra-550b-a55b:free(wasz-ai/glm-4.5-air:free). Theglm,kimi, andnemotronaliases were repointed to their successor models (z-ai/glm-5.2,moonshotai/kimi-k2.7-code,nvidia/nemotron-3-ultra-550b-a55b:free), and the onboarding shortlist was updated to match the new catalog. - OpenRouter app attribution:
HTTP-Referer→https://knightcode.raghavseth.inandX-Title→KnightCode(was "KnightCode CLI").
Removed- Drop two unused dependencies from
@knightcodeai/cli:pretty-ms(never imported) andhono(the toast provider'suseMemonow imports fromreactinstead ofhono/jsx). - Drop discontinued/older version models:
z-ai/glm-4.5-air:free,deepseek/deepseek-v4-flash:free,z-ai/glm-5.1,moonshotai/kimi-k2.6, andnvidia/nemotron-3-super-120b-a12b:free, along with theirglm_airanddeepseekaliases.
FixedTabmode cycle so it reachesAUTO: previouslyTabonly toggled betweenBUILDandPLAN, makingAUTOselectable solely via the/agentsdialog.Tabnow cyclesBUILD → PLAN → AUTO → BUILD.
- Memory follow-ups: feed recent tool usage into the recall selector as an extra relevance signal, frame extraction's "new messages" window from a per-session cursor (so durable facts mentioned during gate-skipped turns are still reconsidered), and drain any in-flight memory extraction on
v0.2.0
Standalone query engine, concurrent tool scheduler, and Apache-2.0 licensing. This release replaces the React `useChat`-based chat harness with a dedicated, framework-agnostic query engine, adds a concurrency-aware tool scheduler, and hardens the interactive terminal experience. The project is now formally licensed under Apache-2.0.
Added- Standalone query engine. A new engine loop drives a turn end-to-end, independent of the React render tree (
lib/engine/). It owns engine event and params types, a transcript-repair pass that resolves dangling/unresolved tool calls, and tool-gating decisions backed by a loop guard to prevent runaway tool cycles. - `useQueryEngine` hook. A thin React hook that drives the engine loop and replaces the previous
useChatharness entirely. - Concurrency-aware tool scheduler. Engine-owned scheduling policy runs tool rounds with bounded concurrency. Introduces an engine
ToolHostcontract and a hook adapter so the engine can execute tools without depending on the UI layer. - Cross-session project memory. Durable, non-obvious facts are extracted automatically after completed turns into a per-project store (
~/.knightcode/projects/<cwd>/memory/) with aMEMORY.mdrecall index. Relevant memories are recalled into the system prompt, a consolidation ("dream") pass merges and prunes the store, and aMemorytool lets the model review, correct, or forget entries. - Per-row tool spinners. Concurrently running tools each get their own inline spinner instead of a single shared indicator.
- `@`-mention path expansion. Paths referenced with
@in a prompt are expanded into the model's context at submit time. - PostToolUse `systemMessage` surfacing. Messages emitted by
PostToolUsehooks are now surfaced to callers. - Apache-2.0 license. Added root
LICENSEandNOTICEfiles andlicensefields in the workspace and CLIpackage.json.
Changed- Extracted
compactHistoryout of the olduse-chatmodule and moved chat message types intolib/engine/messages. - Exposed a hook-free
executeRegisteredToolfor engine use. - Unified all interactive prompts onto a single shared permission panel.
- Dropped the unused
sessionIdfromQueryParams. - Pointed repository URLs at the KnightCodeAI org and scoped the publish workflow to publishable paths.
Fixed- Quit behaviour:
/exitis now the only way to quit; Ctrl+C never exits. - Permissions: every confirm-gated tool now shows a permission prompt, and every awaited tool decision is guaranteed a resolvable prompt; scoped the always-allow sweep correctly.
- Markdown rendering: convert
<br>to real line breaks in prose, expand<br>table cells into continuation rows, and stop rendering literal<br>tags. - Interrupts: render the interrupted marker after the partial response, with a plain interrupted notice (no emoji or completion verb); render interrupted aborts and surface queued mid-turn submits.
- History integrity: stop schema-validating history and instead quarantine invalid tool calls.
- State sync: synchronize message-ref writes, guard submit re-entry, persist the final turn snapshot, queue mid-turn submits, clear finished todos, and only clear the compacting state when it was actually set.
- Hardened file reads, question cancellation, and transcript text handling, plus a sweep of code-review findings across the engine and UI.
- Standalone query engine. A new engine loop drives a turn end-to-end, independent of the React render tree (
v0.1.0
Added- Initial public release:
knightcodeships as a self-contained compiled binary (no Bun required) distributed via platform-specific npm packages, with a headless--versionanddoctor, embedded database migrations, and a non-blocking update check.
- Initial public release: