[v1.20.0] - 2026-09-19
Fixed
- A multi-line
editno longer deadlocks the run (#127 by @brlucasdx). pi already recovers aneditsargument sent as a JSON string instead of an array, and that recovery is a bareJSON.parsein atry. It fails on the single most common way a small local model writes one: raw newlines insideoldText/newText, which is what multi-line code is.JSON.parsethrowsBad control character in string literal, the emptycatchswallows it,editsstays a string, and schema validation refuses the call withedits.0: must be object. brlucasdx counted six in one session. The consequence is worse than a failed call:writeis refused for a file that already exists, so a model that had diagnosed the bug correctly had no way left to deliver the patch at all. This cannot be fixed from an extension (pi's loop runsprepareArgumentsthenvalidateToolArgumentsthenbeforeToolCall, so validation has already rejected the call before anytool_callhook sees it), so it is apatch-pi.mjspatch, applied to both copies of the edit tool pi ships. Control characters inside string literals are escaped and the parse retried; structural whitespace between tokens is left alone, and an escaped\"does not flip the scanner out of the string, which is exactly the case in code being edited. - Behind a router, each model keeps its own context window (#121 by @araujoigor). Two bugs, one line apart. The probe asked
/v1/modelsaboutentry.models[0].id(array position, not the model the user declared as default), and thenwithContextWindowstamped whatever came back onto every model. araujoigor's llama-swap setup has 64k/128k/256k presets of one model with the 64k listed first and the 128k declared default: all three registered as 64k, and the default compacted at half its real window. Nothing was logged, because the probe "succeeded". The window is not a readout: it drives read-guard truncation and the whole context budget. The declareddefaultis now threaded throughmodels.json(splitting on the first/only, since llama-swap preset ids contain slashes), and a router listing is stamped per model from its ownmeta.n_ctx/--ctx-size, with no blanket fallback: a model the router cannot describe keeps the number the user chose rather than another model's measurement. A direct llama.cpp server is unchanged: it serves one model, so its/propsn_ctxreally is every model's window. The swap-time re-probe (#54) now re-stamps only the model being selected, and a selection of a model we never registered is ignored instead of emitting a "context window updated" notice for a change that did not happen. One/v1/modelsfetch now answers all three questions the startup path used to ask separately. - A tool call in prose is no longer "recovered" (#96 by @cal101). Working on JS, cal101 got
harness intervention: the model wrote 1 tool call(s) as text [myCustomAction], followed, delightfully, by the model replying that it had not done anything. It had not. #117 narrowed the nested shape by also requiring an argument key, but a flat object still matched on"name"alone, which is the shape of a config blob or an array of records. The reliable discriminator is not the object's shape, it is whether the tool exists: the nudge asks the model to re-issue the call natively, and a tool that is not registered cannot be issued natively by anyone, so for an unknown name the intervention was not merely a false positive but an instruction the model could not carry out. Names are now checked againstctx.getAllTools()(the full configured set, not the currently active one, since a real tool parked by tool-gating is still a slip worth catching), case-insensitively. A ctx without the accessor filters nothing, since guessing would silently drop real recoveries. LFM2/Liquid calls are exempt: their<|tool_call_start|>tokens appear in no prose by accident, and their--jinjadiagnostic is about the channel rather than the tool. - Compaction no longer arrives late on Ollama (#128 by @brlucasdx). Measured, not inferred: ~50k tokens of Python sent to Ollama at
num_ctx32768 came back withprompt_eval_count: 16386: half the window, plus two, no error, and a constant placed at 50% depth answered "not found". Ollama reports its prompt size after truncating it. The existing silent-overflow check (usage.input > contextWindow) can therefore never fire there, the reading climbs to its ceiling and stops, and by the time the threshold is reached the runtime has been dropping the older half of the conversation on every request. brlucasdx watched the status bar read110.0% / 33k, with the agent re-deriving the same diagnosis and re-reading files it had already read. For Ollama only, the watchdog now measures what is actually in context (sessionManager.buildContextEntries()plus the system prompt, at 4 chars/token) and compacts on the larger of the two figures.maxrather than "prefer the estimate": the estimate omits tool schemas and per-message framing, so it is a lower bound, and taking the larger can only move compaction earlier, never later. Every other provider reports honestly and is untouched;LITTLE_CODER_LOCAL_CONTEXT_ESTIMATE=1/0forces it either way. The first time the gap looks like truncation you get one warning naming both numbers, because the real fix is yournum_ctx. - Skill cards actually load in the pi package (found while reviewing #62).
skill-injectandknowledge-injectresolvedskills/by counting three directories up from the extension, which is right in this checkout (.pi/extensions/<name>/) and wrong in the built pi package (extensions/<name>/). Both loaders end inif (!existsSync(dir)) return;, so the package would have shipped the extensions with none of their content and no way to tell. They search upward for the directory now, which is correct in both layouts.
Added
- Your project's
AGENTS.mdis read again (#104 by @highlyunavailable, confirmed by @dcazrael).--no-context-filesis deliberate (little-coder'sAGENTS.mdshould be the system prompt, not whatever sits in the cwd) but it also threw away the project's own file, which serves a different purpose. highlyunavailable stated the cost plainly: with nothing telling it what the repository is, the model globs the whole tree at the start of every run. dcazrael's is sharper: hisAGENTS.mdexists to say "RULES.md is mandatory, LLM.txt is the map", i.e. to control discovery, so a file that arrives only if the model happens to find it defeats the point. The newproject-contextextension walks up from the launch directory for the nearestAGENTS.md(thenCLAUDE.md) and injects it. It adds rather than replaces, and says so in the block, so a project file cannot quietly re-specify the harness. It is capped at 4000 characters with truncation reported to both you and the model, because the lean cold start is the product. It is injected once, as a hidden tail message rather than a system-prompt append, so the cached prefix survives (#73); the repo's own guard test caught the first draft doing it the expensive way. little-coder's ownAGENTS.mdis never loaded this way, so running it inside its own checkout does not inject a second copy of the prompt it is running on./project-contextshows what loaded;LITTLE_CODER_PROJECT_CONTEXT=0turns it off. little-coder -p --plan-modeworks (#95 by @cal101). A local model is a slow shared resource, so cal101 queues jobs, runs them sequentially in the background, and checks the output later. Plan Mode could not take part: its middle is interactive, and under-pthere is nobody to ask. Of the three policies that make sense (skip the questions, pre-supply the answers, or stop after generating them) this ships the first, the only one that survives an unattended queue. The questions are still generated and are handed to synthesis to be answered from the research, with the plan opening on an Assumptions section stating what was decided for each and why. That costs no extra model call, which matters at 4-7 tok/s, and it puts the assumptions where a batch user can review them rather than in a dialog nobody saw. The plan is printed and written to.pi/approved-plan.md, so the next job can/implementit. The interactive shape could not simply be reused: print mode awaits onesession.prompt()and reads the run's verdict off the last message, so swallowing the input and orchestrating detached tears the process down mid-plan (the #115 hazard). A batch run awaits the research inside the input handler and then lets the original input through, so the ordinary agent turn pi was about to run is the synthesis turn, inside the await print mode is already holding.- Write-capable sub-coders (#93 by @aole).
LITTLE_CODER_SUBCODER_ACCESS=writeaddsedit/writeto a child's tools. Off by default and staying that way: read-only children are what make fanning them out safe, since their answers come back as text and two of them cannot race on the same file, and a child that misunderstands its task wastes tokens rather than the working tree.dispatchis withheld at both levels regardless: a child that can spawn children is a fan-out bomb, and that does not change when it can also write. Thedispatchtool description reports the level in force, since a model told children "CANNOT edit or write files" will not delegate work that needs an edit. pi-little-coder, a pi package (#62 by @dbmrq). The extension and skill layer, published alongside each release for vanilla pi users:npm install pi-little-coder. Extensions are discovered from.pi/extensions/*/index.tsat build time rather than enumerated, so a new one cannot be forgotten: this release'sproject-contextwas picked up automatically, and then deliberately excluded, since pi discovers context files itself and shipping it would inject a second copy. It is a curated subset of the extension layer, not little-coder-in-a-box, and the generated README says so:scripts/patch-pi.mjs, the explicit--no-extensionsload order, the global-settings merge, the llama.cpp context re-probe and the update flow all live in the launcher and do not come along.
Changed
manualmode asks abouteditandwritetoo (#122 by @marek2901). It gated shell commands only, so the tools that actually change your files went through unprompted, which is not what the name promises. They now offer Apply, Deny, or Apply all (this session), and a session with no UI denies, which is the right default for a mode built around a human in the loop and matches the precedent #90 set.write-guardstill runs on top of an Apply. Sub-coders are unaffected: they are pinned toauto.2>/dev/null || trueis no longer refused.trueandfalseare on the whitelist. Neither builtin can do anything, and refusing them was actively misleading: @guppy42 reported the model concluding from the refusal thatlswas the problem and switching toglob.
[v1.19.0] - 2026-08-29
Fixed
little-coder -psurvives a compaction instead of losing the run (#115 by @Or1j1n). v1.18.0 stopped the crash in #108 but not the thing it was a symptom of: under-p, a long task that the TUI completes exitedRequest abortedwith code 1 and empty stdout. pi'scompact()is the manual path and its first act isawait this.abort(), and print-mode reads the run's verdict straight off the last message (if (stopReason === "error" || stopReason === "aborted")), so aborting the run discards the answer. There is no safe way to call it from inside a headless run at all: pi holds_isAgentRunActivetrue for the whole of_runAgentPrompt, soabort() -> waitForIdle()cannot resolve while the run is on the stack. Calling it fromturn_startaborts the run; calling it fromagent_enddeadlocks instead, which is not a guess (a run hung past 400s before that approach was abandoned). pi's own threshold compaction has neither problem, because it runs inside_handlePostAgentRunand ends withreturn this.agent.hasQueuedMessages(), which the caller turns intoagent.continue(). So headless no longer compacts at all: it queues the continuation fromsession_compact, and pi carries the run forward inside print-mode's own await. The mid-run watchdog stays TUI-only, which is where #59 needed it and where Or1j1n confirmed it already works. Measured end to end against Qwen3.6-35B-A3B: before, the answer was lost; after, one compaction, two user messages (the prompt and the queued continuation), exit 0, correct answer.- The compaction resume no longer errors with "Agent is already processing a prompt" (#114 by @guppy42). The resume went out as a bare
sendUserMessage, which routes throughprompt()and throwsAgent is already processing. Specify streamingBehavior ('steer' or 'followUp') to queue the message.whenever the agent has not fully stopped. pi catches that and surfaces it as an error, which is what showed up mid-compaction. It is sent as afollowUpnow, so it queues; and a queued message is exactly what the #115 fix needs, so the two were one change. - A bare-JSON tool call whose arguments are an object is recognised again (#117 by @mth-farias). The recovery pattern was
/\{[^{}]*"name"\s*:\s*"(\w+)"[^{}]*\}/g, and[^{}]*excludes braces outright, so it could never match{"name": "read", "arguments": {"path": "foo.py"}}, the single most common way a model writes a call as text. Because the "re-issue this natively" nudge is driven off this list, the miss was silent: nothing ran and nothing was reported. It now scans brace-balanced objects (the same scanner the #102 fix introduced), and the argument key may beinput,parameters,argumentsorargs, including OpenAI's JSON-string form. The nested shape is only accepted when the object also carries an argument key, so prose containing a config blob with anamefield is still not read as a tool call (#96's failure direction). - A sub-coder tracker can no longer crash the agent after a session is replaced (no issue; found while fixing #119).
SubCoderTrackerholds actxand toucheshasUI/ui.setWidgetfrom an animation timer and fromend()in afinally, and every accessor on an invalidated ctx throws./new,/clearand (since v1.18.0)/implementcan all replace the session while a dispatch is in flight. This is the third instance of one bug shape in a month, after #108 and #119, so it is now a shared helper:_shared/safe-ctx.tsswallows exactly pi's stale-ctx error and rethrows everything else, so a real bug in a UI call still surfaces instead of being eaten. - Router-mode servers no longer hide their own models (#112 by @NoelJacob). Thirteen presets on the server, none of them in
--list-models, and selecting one failed as "model not found", becausemodels.jsonis a curated list and a router serves whatever the user configured. The llamacpp startup probe now also reads/v1/modelsand registers the served ids it finds, each with its own window (meta.n_ctxwhen loaded,--ctx-sizefrom the recorded launch args when not). Discovery only ever adds, never shadows a declared id, and does nothing at all for a single-model listing. That is the ordinary local case, wheremodels.json's friendly alias is the better name and the raw*.ggufid beside it would just be noise. - The llamacpp context probe works behind a router and behind an API key (#116 by @bjornclauw).
/propssits behind--api-keymiddleware and a llama-swap router answers it itself with no usablen_ctx, so the probe failed on every launch and silently fell back to the declared window, which drives the read-guard and context budget, not just the readout. Probes now send the key as a Bearer token and fall back to/v1/models. Verified here against a live-c 131072server:props: 131072,models: 131072. /newwith a background job running no longer kills the process (#119, fixed by @ktutumi in #120).reapAll()sends SIGTERM asynchronously and clears the job map, so a child'sclosecould arrive after pi had disposed the old runtime, and the handler readhasUIoff a stale ctx. Beyond the crash, latedata/closeevents from a reaped job could repaint and wake the replacement session; they are now ignored, whilejob.exitedis still recorded first so the delayed SIGKILL escalation from #102 keeps working.
Added
/skills(#118 by @marouamghar). pi's/skill:nameaddresses pi skills; little-coder's tool skill cards are a different mechanism (selected per turn by error-recovery > recency > intent, injected at the conversation tail), so pi's command cannot see them and there was no way to check what had loaded./skillslists the cards and their token cost,/skills <tool>pins one ahead of every automatic signal for when the selector keeps picking a different card, and/skills offhands selection back.
[v1.18.0] - 2026-08-22
Fixed
- A compaction that settles after the session is gone no longer takes the process down (#108 by @heinrichI). The context watchdog reported its result through the
ctxcaptured at the turn that fired the compaction, and pi invalidates the whole extension runtime ondispose(), soctx.uiandpi.sendUserMessage()throwThis extension ctx is stale after session replacement or reloadfrom that point on. pi invokes the compaction callbacks from a floatingvoid (async () => …)(), which turned that throw into an unhandled rejection and a hardexit 1. heinrichI hit it in a sub-coder, where the headless child disposes the moment the agent settles, so a compaction fired near the end of a run was racing teardown by construction. The UI handle is now captured while the ctx is known-good and every post-compaction use of it is best-effort: nothing the watchdog wants to tell you is worth crashing over. - "Already compacted" no longer halts the run (#109 by @guppy42, #91 by @charly1r, also reported by @cal101). pi's
compact()reports through one promise with two outcomes, and the #68 loop guard treated every rejection as "compaction is futile, pause". Two of the three rejection shapes are not that.Already compactedmeans another compaction landed first and pi'sprepareCompactionfound the branch's last entry is a compaction: the context is compacted, only our call lost the race, so pausing there stranded a mid-task run at the prompt behind a "could not proceed" warning until the user typed "resume".Compaction cancelledmeans the run was aborted or the session is being torn down, which is not a failure to report at all. Outcomes are classified now:Already compactedresumes the run exactly like a successful compaction, a cancellation is silent and leaves the watchdog armed, and only a real failure (Nothing to compact (session too small), a provider error) still pauses. - A turn boundary during an in-flight compaction can no longer start a second one (found while fixing the above; no issue). The
compactingflag was cleared at everybefore_agent_startso a lost callback could not wedge the watchdog off permanently. But pi'scompact()aborts the run and reconnects the agent while it is still summarizing, which makes a genuine in-flight compaction one of those boundaries. Clearing the flag there let a secondcompact()fire on top of the first, and the one that lost is exactly theAlready compactedabove. An outstanding-call counter now means only a stale flag is dropped. - A
/dev/nullredirect butted straight up against a chain operator is no longer read as a file write (#107 by @ashalliants, a regression of #87). The write guard's word splitter broke on whitespace and redirects but not on control operators, sofind … 2>/dev/null; find …parsed its target as/dev/null;(not a path the device exemption knows), and an ordinary two-findone-liner was refused as an unsafe write. ashalliants' diagnosis was exact, and the same hole was open for2>/dev/null&&,||,|, and a closing subshell paren, none of which need a space in front of them.;,&,|,(and)now end a word like>and<already did, with fd duplication (2>&1,>&2,>&-) checked on the raw text so it is still recognised for what it is.
Changed
- Approving a plan saves it;
/implementruns it (#105 by @dcazrael, for #98 by @heinrichI). heinrichI's report was that planning fills the context window and the model then loops once implementation starts on top of it. It does, and the fix has to be a fresh session, which pi exposes only to command handlers (ExtensionCommandContext.newSession), never to theagent_endhandler where Plan Mode asks for approval. So approval now persists the plan to.pi/approved-plan.mdand stops;/implementreads it back, switches to the action model, replaces the session with one seeded with the plan as a hidden context entry, and starts the work. Moving the phase handover from approval to/implementis worth having on its own: on a single llama.cpp backend a handover evicts and reloads weights, so approving a plan you then decide to rewrite used to cost you a reload for nothing. manualpermission mode actually asks (#90 by @marek2901). It blocked exactly whatautoblocks and differed only in the wording of the refusal, which is not what the name promises. It now shows the command and prompts before running it, whitelisted or not: in manual mode the user is the whitelist.write-guardstill runs on top, so a confirmedcat > existing.pyis still caught. A session with no UI (headless) refuses everything in this mode, which is the right default for a mode built around a human in the loop, and sub-coders are unaffected because they are pinned toauto.
[v1.17.0] — 2026-08-16
Added
- Qwen3.8-27B (dense + MTP) in the shipped registry. Its NextN head is in the GGUF (
qwen35.nextn_predict_layers=1,blk.64), so MTP speculative decoding works — measured draft acceptance ~0.87. Measured on an RTX 5070 Laptop (8GB) withUD-Q4_K_XL: 6.42 tok/s at 32k context,-ngl 18, 7042MB; 6.72 tok/s at 16k with-ngl 20. Being dense, it has no experts to park in RAM, so there is no--n-cpu-moetrick and it runs about 7× slower thanQwen3.6-35B-A3Bon the same card — the quality option, not the fast one. One trap worth knowing: at 32k,-ngl 20passes/healthand then generates zero tokens. It fits in VRAM with nothing left to compute, so a server that started is not a config that works. Launch script:run/qwen38-dense.sh. - README sections for background jobs and per-phase model selection, both shipped in v1.16.0 without docs.
[v1.16.0] — 2026-08-16
Added
Background shells that wake you on job events, not on a timer (new).
bashblocks the turn until a command exits, so a long job either freezes the session or gets polled — and polling a six-hour fine-tune every five minutes is 71 wasted turns on a machine where each one costs real seconds.ShellStartruns a command in the background and returns immediately; you declare what is worth interrupting you for and the harness stays silent until it happens:{"name": "ShellStart", "input": {"command": "python train.py", "label": "finetune", "wake_on": {"match": ["Traceback", "CUDA out of memory", "val_loss="], "every_n_matches": 10, "silence": "15m"}}}Wake rules are
exit(default on),match(regex, falling back to literal text),silence(stalled after producing output), andevery_n_matches(throttle a chatty pattern). Urgency picks the delivery lane: a crash or an error-ish match interrupts the current turn, a clean exit or a milestone waits for the tool calls in flight, routine output rides along with the next turn. Wake payloads are bounded — an excerpt plus the exit code, never the whole log — withShellLogto page deeper on demand. AlsoShellList,ShellSend(stdin, for a REPL or a prompting installer), andShellStop. A footer line shows what is running.Jobs may outlive a turn, never the session:
session_shutdownand every catchable signal reap them, and because SIGKILL is catchable by nobody, each job additionally carries a watchdog that kills its own process group when little-coder's pid disappears. Jobs run in their own process group and are signalled as a group, sopython train.pyunder a shell dies with it rather than being orphaned holding the GPU.Per-phase model selection (#61 by @cndjonno, with @cal101). Plan on a big model, implement on a small one.
/plan-modeland/action-modeltag them (with autocomplete and fuzzy matching, so/action-model 9bresolves),/phase-modelsshows the state, andmodels.jsonsupplies defaults viaplanModel/actionModel. The tags are live session state rather than launch config, because the use case that motivated this is swapping planners mid-session to A/B them. Entering Plan Mode switches to the plan model; approving a plan hands over to the action model./model-handover manualturns the automatic switching off — worth knowing that on a single local backend a handover evicts and reloads weights, so "never switch for me" is a performance choice as much as a taste one. Untagged phases use the active model, so an unconfigured session behaves exactly as before.
Fixed
- A refused shell command now says what to do instead of just what was refused (#94). Observed live: a refusal on
./build.shwas followed bybash ./build.sh, thensh ./build.sh, then a successfulpython3 -c "subprocess.run(...)"— three wasted turns and the guard defeated anyway, since interpreters are themselves whitelisted. The refusal now names the evasions not to attempt, points atedit/writefor anything that does not need a shell, and gives the user the actual remedy (LITTLE_CODER_BASH_ALLOW="<cmd>"). Same scenario after the change: refused once, the rest of the task completed, the remedy reported, no bypass attempted. This is the token-burn half of #94; the whitelist's porousness is unchanged and still open there. ShellSessionno longer claims a persistence it does not have. It advertised "cd, env vars, and shell state persist across calls", which is true only under Terminal-Bench's tmux backend. Locally the backend isexecSync— one process per call — so acdappeared to work and silently did not apply to the next call. The description is now computed per backend, and points atShellStartfor anything long-running.- The list of shell-executing tools is shared by both guards (no issue; found while adding
ShellStart).permission-gateandwrite-guardeach kept their own copy, which is exactly how #70 happened — the gate knew aboutbashbut notShellSession, so a refused write simply went through the other tool. One list now, in_shared/shell-write.ts, with a test that fails if a shell tool is gated by one guard and not the other.
[v1.15.0] — 2026-08-15
Fixed
- Tool calls resolve again on pi 0.83 (#92 by @rfairburn, confirmed by @highlyunavailable). pi 0.83 registers its built-ins lowercase (
read,write,edit,bash,grep,ls), but the skill cards andAGENTS.mdstill taughtRead/Write/Bash, so a request to runlscould fireBrowserNavigateinstead. All tool names inskills/tools/*.md,INTENT_MAP, and the system prompt now match what pi actually registers. - The skill selector's error-recovery and recency priorities work again (found while fixing the above; no issue).
skill-injectkeyed its registry bytarget_tool(Bash) but looked it up with the names pi reports on tool events (bash), so two of the selection algorithm's three priorities silently matched nothing from v1.14.0 on. Lookup is case-insensitive now, so a future rename degrades instead of going dark. - Sub-coders are no longer told to use tools they are forbidden to call (#97 by @heinrichI). The evidence-first research protocol was injected into sub-coders, whose allow-list has no
Evidence*tools — and its step 4 ("callEvidenceListbefore answering; if it is empty you are not ready") is then unsatisfiable by construction. Injected guidance is now filtered against the process's real allow-list; children cite inline in their report instead. Thelc.isSubtaskguard that was meant to prevent this had never been set by anything. - A sub-coder that ignores SIGTERM is force-killed (#102 by @ltanon-ai). The SIGKILL escalation was gated on
proc.killed, which Node sets when a signal is dispatched rather than when the process exits, making the branch unreachable and leaking hung children as orphans. - The pre-edit checkpoint backup actually runs (#102). It keyed on
input.file_path, but pi'swrite/editpasspath, so~/.little-coder/checkpoints/had silently never been written for normal operation. - A valid tool call followed by trailing text is no longer dropped (#102). The output parser's last-resort
/\{[^{}]*\}/fallback grabbed the first brace-free object, returning a nestedinputfragment with noname. Replaced with a quote- and nesting-aware scanner. - The update check validates the version it gets from the registry (#102). A non-semver
latestreachedcompareSemver(NaN math) and, on Windows,cmd.exe /c npm install little-coder@<latest>. Also keeps the full prerelease tail (1.0.0-rc-1) instead of truncating at the first hyphen. package-lock.jsonpins its dependencies again (#103 by @Thib-ai). Three nested@earendil-worksentries shipped without anintegrityhash, sonpm install -g little-coderfetched whatever the registry served rather than verifying the tarball, and the repo could not be packaged for NixOS. A CI job now runsnpm cion every PR so it cannot regress.
Changed
AGENTS.mdno longer promises "full system access" (#94 by @Franck-Nein, with @mac-edmondson, @colinmutter, @steverhoades and @guppy42). colinmutter's diagnosis was right: the prompt told the model it had unrestricted access while the shell guard told it otherwise, and the model resolved the contradiction by hunting for a way around the guard — burning a lot of tokens doing it. The prompt now states that a refused command is an answer, names the specific evasions not to attempt, and points atedit/writefor anything a shell isn't needed for. The autonomy framing is unchanged. This addresses the token burn, not the underlying porousness of the whitelist; see the issue for that.
[v1.14.0] — 2026-07-31
Changed
- Bundled pi upgraded 0.79.4 → 0.83.0 (#85 by @m00kfu). Modern third-party extensions that import newer pi-ai paths (e.g.
pi-ai/…/compat) now resolve instead of failing to load. Our runtime patch was re-based onto pi 0.83 and the full test suite re-verified against it.
Fixed
/dev/nullredirects are no longer blocked (#87 by @manueloverride). The shell-write guard counted2>/dev/nullas a destructive file write; it now exempts the null-ish character devices (/dev/null,/dev/stdout,/dev/stderr,/dev/tty,/dev/zero,/dev/random,/dev/fd/N). A redirect to a real file alongside them is still caught.- A backend that rejects the thinking level no longer loops on "empty response" (#86 by @aole). A provider 400 (e.g. an ollama model that doesn't support thinking) is now treated as an error turn, not an empty model response, so it isn't re-sent three times over pi's real error — which is shown, with a one-line hint to lower the thinking level.
[v1.13.0] — 2026-07-30
Added
--plan-modestarts a session already in Plan Mode (#84 by @aole).little-coder --plan-mode(orLITTLE_CODER_PLAN_MODE=1) opens with the◆ PLAN MODEindicator on;ctrl+qstill toggles. Interactive sessions only — headless and sub-coder runs ignore it.
Fixed
- Re-running a build command after editing a file is no longer mis-flagged as a loop (#81 by @manueloverride, workaround by @Franck-Nein). The quality monitor now skips its "repeated tool call" verdict when the previous turn also ran a state-changing tool (Edit/Write/Bash), since the environment changed between the two identical calls. A verbatim repeat with nothing else changing is still caught.
[v1.12.0] — 2026-07-24
Fixed
little-coder no longer ships an npm install script, so Socket's malware scanner has nothing to flag (#75 by @modemlooper, diagnosed as a false positive by @abhisek).
npm install -g little-coderwas producingPotential malware detected with AI scanfrom Socket Firewall, pointing at thepostinstallhook. The hook ranscripts/patch-pi.mjs— a visible, dependency-free file in the repo that re-applies two cosmetic source edits to the bundled pi — so the alert was a false positive on the shape of the code (an install script touchingnode_modules) rather than on anything it did. Rather than just explain that, v1.12.0 removes thepostinstallentry entirely: the launcher already callsapplyPiPatches()on every launch, and every self-updating user was skipping the postinstall anyway because/updateand the launcher's auto-update both install with--ignore-scripts(#50). Launch-time patching is now the only path, which also means it self-heals if pi is reinstalled underneath.scripts/patch-pi.mjsitself is unchanged in behavior and still shipped — it's imported by the launcher, no longer executed by npm.The launcher's pi patching and pi-changelog suppression were both silently dead (found while investigating the above; no issue).
bin/little-coder.mjsresolved pi's package root inside afor (const piPkgRoot of piPkgCandidates)loop and then referencedpiPkgRoottwice further down — at the patch call and at thelastChangelogVersionpin. Afor (const …)binding is scoped to the loop body, so both references threwReferenceError, and because both sit inside deliberately best-efforttry/catchblocks the failure was completely invisible. Two real consequences: our pi runtime patches never re-applied at launch (so a--ignore-scriptsupgrade left pi unpatched), andlastChangelogVersionwas never written — which meant pi's own upstream "What's New" changelog rendered inside the little-coder TUI after every bundled-pi bump. The binding is now declared at module scope and assigned on the successful candidate. Two end-to-end regression tests (bin/launcher-pi-root.test.mjs) run the real launcher and assert both effects land; both fail against the old code.ShellSessionno longer routes around the write guard and the permission gate (#70 by @rvanswieten). Qwen3.6-35B-A3B, refused a whole-filewrite, switched tocat > backend/main.py << 'ENDOFFILE'and got the same bytes anyway — 5 times in one session (main.py×3,App.jsx,pyproject.toml). rvanswieten's diagnosis was correct on all three counts:write-guardmatched onlytoolName === "write",permission-gatematched onlybash/Bash, andShellSessionhit neither, landing straight inexecSync. The bash whitelist wouldn't have saved it either —isSafeBashwasstartsWithon the raw string andcatis whitelisted, with no awareness of the>immediately after it. Fixed with a shared, pure command analyzer (_shared/shell-write.ts) that strips heredoc bodies (so a>or an apostrophe in the payload can't confuse it) and reports the paths a command writes to via>,>>,tee, ordd of=, ignoring fd duplication (2>&1,>&2), process substitution, input redirects, and anything inside quotes. Three behavior changes follow:ShellSessionis gated exactly likebash; every command in a&&/||/;/|chain must pass the whitelist independently (ls && rm -rf /is refused on therm, not admitted on thels); and any command that writes through the shell is refused inauto/manualmode, with a message naming the path and pointing atWrite/Edit. Inaccept-allmode (benchmark runs) the whitelist is still skipped, butwrite-guardnow inspects shell writes too, so a redirect that would clobber an existing file gets the same Edit recipe thewritetool gives, and a reserved Windows device name (#60) is refused even as an append. A>>append to an existing file is allowed — nothing is destroyed, so it isn't the whole-file-rewrite failure mode the guard exists to prevent. Deliberately not matched on the heredoc delimiter string, which any other delimiter would have defeated.Per-turn skill and knowledge injection no longer destroys the KV cache (#73 by @manueloverride, with @charly1r). manueloverride caught this with
cache-hunter: llama.cpp suddenly re-churning 120k of message history "for no reason" mid-conversation, because the harness had changed the beginning of the prompt. Exactly right.skill-injectandknowledge-injectappended their selected blocks to the system prompt, which is the first thing in every request — and since both blocks are recomputed per turn from the user's prompt, they changed on most turns and invalidated the entire cached prefix each time. The fix uses a hook pi already had:before_agent_startmay return amessageinstead of asystemPrompt, which pi appends after the user's message and converts to auser-role message on the way to the provider. The guidance now lands at the tail of the conversation with every preceding byte untouched, so the prefix stays cached and only the new tokens are processed. The recency argument that put these blocks last in the system prompt gets stronger rather than weaker — the conversation tail is as late as placement gets. All four injectors moved onto one shared helper (_shared/inject.ts):skill-inject,knowledge-inject,plan-mode's synthesis instructions, anddeep-research's report brief (the largest of the four, and the one where a system-prompt rewrite cost the most). Blocks are hidden from the transcript (display: false), and a block identical to the previous turn's is not re-sent — the earlier copy is still in the conversation, so repeating it would only spend context.LITTLE_CODER_INJECT_MODE=systemrestores the old placement, which is what the whitepaper scaffold reproduction was measured against. A regression fence now fails the build if any extension returns a rewritten system prompt directly, since the symptom is invisible from inside little-coder.Measured against a live llama.cpp server (Qwen3.6-35B-A3B, 4 turns, same prompts in both modes, counting the prompt tokens llama.cpp actually evaluated per request):
Turn context old: system prompt new: tail message 1 ~10.6k 6,525 6,529 2 ~16.5k 12,311 6,235 3 ~23k 18,488 6,377 4 ~29k 24,586 6,438 total 61,910 25,579 The old numbers climb with the conversation — each turn re-evaluates almost everything, which at manueloverride's 120k context is the 120k re-churn he reported. The new numbers are flat: only the genuinely new tokens are evaluated, whatever the history length. 58.7% fewer prompt tokens over four turns, and the gap widens with every additional turn.
Note for anyone who arrived from the same thread: the llama.cpp/Qwen KV-cache bug charly1r linked is a genuinely separate problem, upstream of little-coder.
The startup header no longer advertises a key that does nothing (#74 by @heinrichI). The header's hint row listed
ctrl-r more, but pi binds "expand / more" toapp.tools.expand=ctrl+o;ctrl+ris bound only inside the session-tree overlay, so at the prompt it genuinely did nothing. Corrected toctrl-o, andf2(deep research) was missing from both the header and thectrl+hshortcuts panel — the one flow people ask about was absent from the panel whose whole job is discoverability. Both now list it, andbuildHeaderis covered by tests that assert every advertised key is a real binding. Thectrl+ohalf of the report is expected behavior rather than a bug, now documented: during a Deep Research run the research sub-coders are separate child processes, so their tool output never enters this session's transcript and there is nothing forctrl+oto expand — the progress bar is the view of that work — and while the max-agents or clarifying-question dialogs are open, keys belong to the dialog.Two panels were silently losing their last rows to pi's widget height cap (found while verifying this release against a live TUI; no issue). pi renders a string-array widget through
content.slice(0, MAX_WIDGET_LINES)— the cap is 10 — and appends... (widget truncated), so anything past ten lines is dropped from the end. Two places were over it: thectrl+hshortcuts panel had grown to eleven rows plus a header, which quietly ate/hotkeys— the row whose entire job is pointing at the authoritative keybinding reference — so it now lays out in two columns and shows every shortcut in seven lines; and the sub-coder tracker appends a row per sub-coder inbegin()and never resets, so a session with severaldispatchturns accumulated rows without bound and pi's truncation then hid the sub-coders that were actually running behind a backlog of finished ones — exactly backwards for a live progress view. The tracker now always shows every running sub-coder, fills the remaining rows with the most recently finished, and accounts for the rest on one… +N earlier sub-codersline, with the header still reporting the true total. Both are covered by tests that assert the ten-line ceiling.
Added
- A first-class, opt-in way to add your own extensions (#67 by @thegausiantheory; #69 by @johnzan, with @charly1r). little-coder launches pi with
--no-extensionsand loads exactly its bundled set, which is what keeps the cold-start context near 7k tokens and makes behavior predictable — charly1r's write-up on #69 explains the tradeoff better than the docs did. But "I want to add my own" is a fair ask, and the only answer was an env var nobody found (LITTLE_CODER_EXTRA_EXTENSIONS) or forking the installed package (#46). Three additions, all opt-in, all no-ops when unused:- A user extension directory.
~/.config/little-coder/extensions/(or$XDG_CONFIG_HOME/…, orLITTLE_CODER_EXTENSIONS_DIR). Each direct child is one extension — a.ts/.js/.mjsfile, or a directory with anindex.ts/index.js. Loaded after the bundled set, so yours can override ours. The directory is never created for you and nothing is written to it; an install that ignores it behaves exactly as before. Survivesnpm install -g little-coder@latest. /extensions— a panel showing what's loaded, split by where it came from (bundled / yours /LITTLE_CODER_EXTRA_EXTENSIONS), whether pi's own discovery is on, and anything that failed to load. Because little-coder forcesquietStartupto keep launch clean,--verboseused to be the only window into this, and a user extension that didn't resolve warned once on stderr moments before the TUI painted over it — those warnings now also fire as a notification at session start. Rendered as a widget rather than a chat message on purpose: a custom message would be sent to the model, spending context on a diagnostic it has no use for.--with-pi-extensions(orLITTLE_CODER_PI_EXTENSIONS=1) — omits--no-extensionsso pi discovers its own extensions from~/.pi/agent/extensionsand./.pi/extensions. Off by default and it prints a notice when on, because the guarantee it gives up is real: the extension set is no longer fixed, cold-start context grows, and a cloned repository can contribute extensions from its.pi/directory (pi's own trust prompt still gates that code).
- A user extension directory.
- New guide:
docs/extensions.md— how to write an extension, where to put it, which bundled extension to read as a reference for each kind of change, the two things to know before writing one (don't rewrite the system prompt per turn; cap rendered lines to the terminal width), and a community list: BMorgan1296's Zed/pi-acpbridge (#58), johnzan's Telegram bridge (#69), and charly1r's llama.cpp configs and local-model tournament results (#63). Community projects are linked, not vendored.
Docs
- The status line is documented (#71 by @johnzan, first answered by @charly1r). A new README section explains every field, read off pi's
footer.jsrather than inferred. Three corrections to the reasonable-looking guess in the thread: the first number is cumulative session input tokens, not current context size;CHis the cache-hit rate of the latest response alone (cacheRead / (input + cacheRead + cacheWrite)), not a session average, which is why it swings; and(auto)specifically means automatic compaction is enabled. A lowCHon a long conversation is a real signal — it's exactly what #73 above was causing. - Three stale README facts corrected. Plan Mode is
ctrl+q, notalt+p(it moved in v1.9.0 and the README never followed). Sub-coder concurrency defaults to 1, serial, not 2 (changed in v1.9.10 by #57). And the claim that pi themes don't load was wrong:--no-extensionsgates extensions only — pi's theme discovery is governed by a separate--no-themesflag that little-coder never passes, so pi themes in~/.pi/agent/themesand./.pi/themeshave always worked. That correction also applies to my earlier answer on #67. - New Troubleshooting entries for the malware alert,
ctrl+r, pi extensions and themes, extension load failures, and the llama.cpp reprocessing symptom.
[v1.11.0] — 2026-07-18
Fixed
- The context watchdog no longer wedges a session into an unrecoverable "Nothing to compact" loop after a mid-run compaction (#68 by @charly1r, reproduced by @Qm-jmz). Symptom: after a mid-run compaction the resumed run started re-reading project files, context climbed back over the threshold, and a second compaction fired — but the only summarizable slice left was the tiny post-compaction tail, so pi threw
Compaction failed: Nothing to compactand, because pi'scompact()aborts and disconnects the agent before that check, the session was left dead (nothing the user typed recovered it). The deep fix (detect the incompressible tail and elide-rescue before giving up) belongs in the pi runtime little-coder wraps and has been flagged upstream; this release adds a wrapper-side guard that breaks the loop before the session can wedge. Thecontext-watchdogextension now measures each compaction's actual effect: if usage is still withinMIN_PROGRESS(5 points) of the threshold afterward — i.e. compaction freed too little — it pauses automatic compaction and tells the user (context still at N% after compaction — automatic compaction paused to avoid a loop) instead of firing a doomed secondcompact(), and it re-arms automatically once usage drops back below the threshold band (a manual/clear,/compact, or a smaller turn). A compaction that errors (onError, e.g. a "Nothing to compact" that still slips through) pauses the same way rather than silently retrying. Separately, the mid-run resume message now explicitly tells the model not to re-read files it has already read — the re-scan is what re-inflated context into that second compaction in the first place. Behaviour is unchanged when the watchdog is disabled (LITTLE_CODER_NO_COMPACT_WATCHDOG=1/LITTLE_CODER_COMPACT_AT_PERCENTout of band). The deep incompressible-tail recovery has been flagged for pi upstream.
Added
- A configurable default model, so bare
little-coder"just works" (#65 by @cndjonno).models.jsonnow takes a top-level"default": "provider/id"(shipped default:llamacpp/qwen3.6-35b-a3b). On launch, if you didn't pass your own--modeland pi has no persisted model selection yet, little-coder injects that default and prints the model's friendly name —▸ default model: Qwen3.6-35B-A3B (MoE, local llama.cpp) (llamacpp/qwen3.6-35b-a3b)— so a first run no longer needs the verbose--model provider/id. It's first-run-only by design: the moment you pick a model in-session (pi persistsdefaultProvider/defaultModel), that choice wins on every later launch and the default never overrides it; passing--model, or a headless/sub-coder run, also skips injection. A user override file's own top-leveldefaulttakes precedence over the shipped one. The friendly name is surfaced at launch rather than swapped into pi's width-constrained footer, which keeps the compact model id there and avoids the terminal-width overflow class from #51. - Auto-relaunch after an in-place update (#66 by @cndjonno). Answering
Yto the launcher'sUpdate now? [Y/n]prompt used to install the new version and then drop you back to the shell with "re-run little-coder to use it." The launcher now re-execs the freshly-installed version in place with your original arguments (adding--no-update-checkso the child doesn't re-poll and can't loop), landing you straight in the new build — it printsRelaunching little-coder…and, if the re-exec somehow can't start, falls back to the old manual-relaunch hint. The in-app/updatecommand still shuts the session down for a clean manual restart, since it runs inside the pi child where an in-place re-exec isn't safe.
[v1.10.0] — 2026-07-06
Added
- Deep Research — a Scope → Research → Write flow that fans out read-only research sub-coders and writes one cited report. Toggle it with
f2(an indicator appears below the input; the next prompt becomes the research topic) or run/deep-research <topic>directly. The flow mirrors the orchestrator-worker pattern: it asks a max-agents cap M (1–10), scopes the topic into 1–4 clarifying questions → a compressed research brief (the north star), has a lead decompose the brief into parallel subtopics, dispatches read-only research agents shown as a rectangle progress bar (not the per-child subagent view), runs a gap-analysis pass that spawns 1–2 more agents only if coverage genuinely falls short, and finally has the main agent write one cohesive, inline-cited markdown report saved todeep-research-<slug>-<timestamp>.md. The agent tree scales with M (e.g. M=6 → 1 lead + 3 wave-1 + ≤2 gap-fill);f2is unbound by pi's app keybindings, the emacs-style editor, and every other extension, so it registers with no conflict. - Project-safety by construction. Research agents get the full read + browse toolset minus
bash— with bash they were scaffolding and compiling throwaway projects (cargo new,go build) in the working tree, once leaving ~200 MB behind — and every reasoning/research child runs in an ephemeral scratch cwd (mkdtemp, removed on exit), so a research run can never write into the user's repo. The one-shot write turn additionally blocksedit/write/bashso the report is emitted as text, saved by the extension itself. - Grounding and honesty enforced in the prompts. Research children must cite the full source URL for every factual claim and end with a
Sources:list, and are told an honest "not found" beats an invented fact; the writer may use only what the findings support, marks unsupported claims rather than fabricating, and lists only real URLs. A per-child watchdog (reasoning 4 min, research 10 min) kills a hung agent (e.g. a browser wedged on a page) so it can't stall the whole run, with a single automatic retry on timeout to recover a transient hang before that subtopic is lost; when an agent is still lost, the writer is instructed to add a short "Research coverage" note disclosing the gap instead of quietly shipping a thinner report. - Validated headless at scale before shipping. The Scope → Research pipeline was extracted into a UI-agnostic engine (
pipeline.ts) so a gated batch harness drives the same code the interactive flow runs, across 25 full research pipelines (10 topics, then 15, at M=6) scored by mechanical heuristics + a local-model judge + hand review. The grounding work is the headline result: real, resolvable sources went from 0 per report to ~28, structure from 82→94/100, with 25/25 substantial reports and no crashes. The interactive TUI layer (dialogs, live progress bar, the main-agent write turn, ESC/abort, plan-mode coexistence) remains a manual-test surface. Tunable viaLITTLE_CODER_DEEP_RESEARCH_MAXand.pi/settings.json(little_coder.deep_research.default_max_subagents, default 10); research fan-out honorsLITTLE_CODER_SUBCODER_CONCURRENCY(default serial) like the rest of the sub-coder machinery.
[v1.9.13] — 2026-07-05
Fixed
- The mid-run compaction watchdog now resumes the task instead of stranding it at the prompt (#59 follow-up by @charly1r). v1.9.12's
context-watchdogcorrectly fired compaction at 80% mid-run — charly1r confirmed the trigger works across models — but afterward the run just sat idle until he typed something, with llama.cpp still showing activity. Root cause: pi's publiccompact()is the manual path (_disconnectFromAgent→abort→ summarize → reconnect, then idle), and pi's threshold compaction is deliberately "compact, no auto-retry — the user continues manually." So the watchdog aborted the autonomous run to compact and never told it to carry on. The fix passes anonCompletecallback tocompact()and, once the agent is reconnected and idle, callspi.sendUserMessage(...)(which always triggers a turn) with a short instruction to continue the task from where it left off on the freshly-compacted context — so a long run now compacts and keeps going, untouched. The in-flight guard still prevents stacked compactions and re-arms on completion. Behaviour is unchanged when the watchdog is disabled (LITTLE_CODER_NO_COMPACT_WATCHDOG=1/LITTLE_CODER_COMPACT_AT_PERCENTout of band).
[v1.9.12] — 2026-07-04
Fixed
- Long autonomous runs now compact before they overflow the context window (#59 by @charly1r). pi only re-evaluates auto-compaction at a user-turn boundary — its check runs after
agent.prompt()fully returns, i.e. once the model stops requesting tools and goes idle. During one long autonomous run that boundary is never reached: little-coder's small models routinely chain dozens of tool-call turns before yielding, so context climbs unchecked and pi only reacts to the overflow error after the fact. charly1r reproduced it precisely — context growing 34k → 40k → … → 64k across many turns with no compaction until the request overflowed a 64k window. A newcontext-watchdogextension closes the gap: it reads live usage via pi'sgetContextUsage()at every turn boundary and, once usage crosses 80% of the window, calls pi'scompact()mid-run — so a single long run compacts at roughly the same point pi would have if the model had paused. Tunable viaLITTLE_CODER_COMPACT_AT_PERCENT(percent; e.g.70to compact earlier);≤0/≥100orLITTLE_CODER_NO_COMPACT_WATCHDOG=1disable it and defer entirely to pi's end-of-run/overflow paths. It's complementary to pi's own compaction (an in-flight guard prevents double-firing) and independent of thereserveTokens/keepRecentTokensknobs, which still govern how much is summarized vs. kept verbatim.
Added
- The launcher's update prompt auto-continues instead of blocking, plus an in-app notice and
/updatecommand (#64 by @cndjonno). When a newer version was published, the launcher'sUpdate now? [Y/n]prompt blocked startup indefinitely waiting on input — an unattended terminal never got past it. It now auto-continues without updating after 10 s (configurable viaLITTLE_CODER_UPDATE_PROMPT_TIMEOUT=<seconds>;0/off/neverrestores the old wait-forever behavior), and the prompt shows the countdown. Two follow-ups from the same request: (2) if you dismiss or time out of the launcher prompt, a one-line "update available" notice now appears inside the running TUI so the pending update isn't lost, and (3) a new/updatecommand installs the latest little-coder (with--ignore-scripts, matching the launcher's supply-chain posture from #50) and cleanly ends the session so you can relaunch into it — no quitting to remember the npm incantation.
Docs
- Guide for running little-coder inside Zed via an ACP bridge (#58 by @BMorgan1296, with @charly1r). little-coder still ships no ACP server of its own —
--mode rpcis pi's internal extension-UI RPC, not the Agent Client Protocol — but the communitypi-acpbridge drives it well: point pi-acp'sPI_ACP_PI_COMMANDat thelittle-coderbinary and every bundled extension/skill comes along. Newdocs/zed-acp.mdwrites up the full setup (Zedagent_serversconfig + a wrapper script that starts/stopsllama-server), generalized from BMorgan1296's working recipe. Marked explicitly as community/unofficial — a first-class ACP transport still belongs in pi upstream, where both projects would benefit.
[v1.9.11] — 2026-06-28
Fixed
- The context window now re-probes when llama-swap swaps the loaded model (#54 by @cndjonno).
llama-cpp-providerprobed the server's liven_ctxvia/propsonly once at startup, so when llama-swap swapped a model under the same endpoint little-coder kept reporting the old window — and that's not cosmetic: the TUI readout, the read-guard, and the context-budget math all follow the registered window, so a stale value mis-sizes the budget. The extension now hooks pi'smodel_selectevent, re-probes/propswhenever the active model changes to a llamacpp model, and re-registers the provider with the fresh window — with a one-linecontext window updated 32k → 128knotice so a drop like 128k → 16k can't silently mis-size things mid-task. It skips the initial selection (startup already probed), no-ops when the window is unchanged or the probe fails, and honors the existingLITTLE_CODER_NO_CTX_PROBE=1opt-out. Per-phase model selection (a big model for planning, a small one for implementation) is tracked separately in #61.
[v1.9.10] — 2026-06-28
Fixed
- Sub-coder concurrency now defaults to 1 (serial), and
=0is honored instead of silently ignored (#57 by @whateverforever, with @charly1r). On a small local setup two sub-coders contend for the same single model server and run slower than one at a time, so 2 was the wrong default — it's now 1, and parallelism is opt-in viaLITTLE_CODER_SUBCODER_CONCURRENCY=2+. Separately,defaultConcurrency()gated onn > 0, so a user who setLITTLE_CODER_SUBCODER_CONCURRENCY=0to force serial execution silently fell back to the default (2) instead. Explicit values are now clamped to a floor of 1, so0and negatives mean "serial". Both the dispatch tool and Plan Mode's research fan-out go through this one function, so Plan Mode now respects the env var too (it could generate up to 4 exploration tasks but now executes them one at a time under the default). - Write is refused for Windows reserved device names (
nul,con,com1–com9,lpt1–lpt9,aux,prn) (#60 by @charly1r). A model treatingnullike/dev/nulland writing to it created a literalnulfile on Windows — backed by a reserved DOS device name, it's notoriously hard to delete.write-guardnow blocks any write whose basename (case-insensitive, extension ignored) is a reserved device name and tells the model to pick a real filename or not write at all. Enforced on every platform, since a literalnul/confile is a mistake everywhere and a landmine the moment a POSIX-authored repo is cloned on Windows. - Launcher now finds the bundled pi under bun's flat global layout (#56 by @kode54).
bun add -ghoists dependencies flat as siblings of the package (…/@earendil-works/pi-coding-agent) rather than nesting them underlittle-coder/node_modules/, so the launcher's hardcoded nested path failed andlittle-coderwouldn't start. It now tries the npm-nested path first, then the bun/flat sibling path, and the error message lists every location it checked.
Added
ctrl-htoggles an on-screen keyboard-shortcuts panel (#55 by @cndjonno). New hotkeys keep getting added (Plan Mode'sctrl-q, the thinking-level cycle, …) and weren't discoverable;ctrl-hnow shows a compact, width-safe list of the keys worth knowing right below the input, and actrl-h keyshint joins the startup shortcut row.ctrl-his genuinely unbound by both pi and the emacs-style editor (and not in pi's non-overridable set), so it registers with no conflict diagnostic. Because little-coder's custom shortcuts are registered with descriptions, bothctrl-q(plan) andctrl-h(this panel) also appear automatically in pi's built-in/hotkeysreference.
[v1.9.9] — 2026-06-22
Fixed
- Plan-mode toggle moved from
ctrl+ytoctrl+qto clear a built-in shortcut conflict.ctrl+yis the editor's built-in yank/paste (tui.editor.yank), so v1.9.8 logged an[Extension issues]conflict diagnostic at startup and the toggle overrode the editor's paste. The emacs-style editor claims nearly every otherctrl+<letter>(line motion, word/line deletes, etc.);ctrl+qis genuinely unbound, and because pi runs the terminal in raw mode (flow control disabled) it arrives as a clean\x11byte on every terminal — so the toggle works without a conflict and without shadowing any editor key. The indicator, leave-mode hint, and the startup shortcut-row CTA now readctrl-q.
[v1.9.8] — 2026-06-22
Fixed
- Plan-mode toggle is now
ctrl+yinstead ofalt+p. Many terminals (notably macOS) deliver Alt+P as the literalπcharacter rather than anESC psequence, so pi's key matcher never matched and the toggle silently failed — pressing it just typedπinto the input.ctrl+yis delivered as a clean control byte, is left unbound by both pi and the editor (no yank handler), and is dispatched to extension shortcuts before any editor handling, so it fires reliably. The plan-mode indicator and "leave plan mode" hint were updated to readctrl-yto match.
Added
- A
ctrl-y planhint in the startup shortcut row so the plan-mode toggle is discoverable alongside the existingesc///ctrl-rhints.
[v1.9.7] — 2026-06-19
Security
- Auto-updater now passes
--ignore-scriptsto npm so a compromised package can't run arbitrary code during upgrade (#50 by @steverhoades). Lifecycle scripts (preinstall/install/postinstall) are the entry vector that Shai Hulud-style worms (and any other npm postinstall malware) use to land code execution the moment a compromised version of little-coder or one of its transitive deps is published. The launcher's auto-update path (bin/update-check.mjs) now invokesnpm install -g --ignore-scripts little-coder@<latest>, matching pi's posture upstream. The notice-only message shown on non-TTY pipelines was updated in the same change so the manual recovery command surfaces the flag too. Scoped to the auto-updater only — first-install viainstall.sh/npm install -g little-coderstill runs scripts so playwright's chromium download (used by browser-extract-retention's live integration test) lands during onboarding; on patch upgrades the binary is already on disk from that first install, so--ignore-scriptsis the safer default there. Two new tests pin the flag in the source (a removal in either the spawn args or the user-visible command string fails the suite).
Docs
- Troubleshooting entries for
--updateand the Windows ≤ v1.9.5 bootstrap caveat (PR #53 by @i-snyder). Documents thatlittle-coder --updateforces an immediate version check bypassing the 12h cache (and the flag is stripped before pi sees argv), and that users on the broken v1.9.5 Windows updater need a one-timenpm install -g little-coder@latestto reach v1.9.6 (after which auto-update works normally). i-snyder's follow-up to PR #52, exactly as requested in the close comment.
[v1.9.6] — 2026-06-18
Fixed
- Auto-update silently failed on Windows (PR #52 by @i-snyder). On Windows
npmisnpm.cmd(a batch-file shim), andspawnSync("npm", …)withoutshell: truereturnsENOENTbefore npm ever launches — the user saw✗ Update failed (npm exit null). Continuing with v1.9.x., where thenullexit code was the tell that npm never ran. The launcher now invokesnpmviaprocess.env.COMSPEC /c npmon Windows (shell: truewould also work but triggers Node 24+'sDEP0190deprecation warning; COMSPEC doesn't). Cross-platform behavior unchanged: POSIX still uses plainspawnSync("npm", …). The failure message is also fixed — when the spawn itself fails (result.erroris set), the launcher now surfacesresult.error.code(e.g.ENOENT) instead ofresult.status(which isnulland meaningless), so users diagnosing future spawn failures get an actionable code instead ofnpm exit null.
Added
little-coder --updateforces a fresh update check (PR #52 by @i-snyder). The launcher caches the registry "latest" lookup for 12 hours;--updatebypasses that cache and fetches fresh from npm, then either updates or prints✓ little-coder is already up to date (v<x>). The flag is stripped from argv before forwarding to pi, solittle-coder --updateno longer errors withUnknown option: --update. Two newshouldSkiptests document the interaction (the flag forces a check; the notice-only mode still applies on non-TTY).
Notes for upgraders
- No CLI-flag or public-API breakage. Windows users on v1.9.5 or earlier should manually
npm install -g [email protected]this once; future updates will work via the in-app prompt orlittle-coder --update.
[v1.9.5] — 2026-06-18
Changed
- Dispatch tool-result panel now word-wraps wide report lines instead of truncating them (PR #49 by @steverhoades, closes #48 and #51). v1.9.4 fixed the width-overflow crash by truncating each panel line to
width - 2with an ellipsis; v1.9.5 replaces the truncation with word-wrap so the full sentence survives across multiple visual lines — a strictly better UX for markdown sub-coder reports than dropping the tail at char 131. The cherry-picked commit (steverhoades's authorship preserved) keeps the wrap helpers (ANSI-aware prefix extraction, long-token chunking for whitespace-free URLs/paths/base64 that would otherwise defeat word-wrap, plain-text word-wrap), and themakeComponent.render(width)is rebased onto v1.9.4'swidth - 2safety margin so wide-unicode chars our char-countvisibleWidthundercounts still can't sneak past pi's strict line-width check. Inspiration for the long-token sanitizer credited in-source to openclaw-cn's tui-formatters.ts.issue-51-repro.test.tsupdated for wrap semantics (4 cases): no emitted line exceeds; the wrapped lines round-trip to the original 134-char sentence verbatim (no data loss); narrow terminal (40 cols) survives; 200-char URL-ish tokens get chunked so wrapping has room to split.
Notes for upgraders
- No CLI-flag or public-API changes. If you upgraded from v1.9.3 → v1.9.4 → v1.9.5, the user-visible difference between the last two is just wrap-vs-truncate in the dispatch tool's expanded report panel — both eliminate the crash. If you saw an ellipsis at the right edge of a sub-coder report on v1.9.4, you'll now see the full sentence wrapped onto the next line instead.
[v1.9.4] — 2026-06-18
Fixed
- Dispatch tool-result panel overflows the terminal on wide report lines (#51, reopen of #48). v1.9.2 capped every line the live sub-coder tracker emitted, but the dispatch tool's result renderer (
subagent/index.ts'smakeComponent) was still ignoring thewidtharg pi passes torender(width)— it returned the precomputed lines verbatim. pi paints the tool-result panel with a 1-char background-color left margin, so any sub-coder report sentence wider thanterminal_width - 1overflowed pi-tui. Crash log line 453 was a 134-char markdown sentence rendered at terminal width 133 → 135 > 133. The same path runs on--resume(pi re-paints saved tool results from session history), so v1.9.2 users still hit it after upgrading whenever they resumed a session with a wide dispatch report saved — that's why @steverhoades caught the regression.makeComponentnow truncates every emitted line towidth - 2using the existing_shared/width.tsutility (2-char safety margin for wide unicode under our char-count-basedvisibleWidthapproximation), so the dispatch panel can no longer crash a session — live, on resume, or anywhere else. Newsubagent/issue-51-repro.test.tsdrivesmakeComponentwith the user's exact 134-char content shape at width 133 and asserts no emitted line exceeds, plus a narrow-terminal (40-col) survival check.
Notes for upgraders
- No CLI-flag or public-API changes. If you saw
Rendered line N exceeds terminal widthon v1.9.2 / 1.9.3 — especially while resuming a session — 1.9.4 fixes it. If you still see it after upgrading, the offending line in~/.pi/agent/pi-crash.logshould let us spot the source; reopen #51 or #48 with the log attached.
[v1.9.3] — 2026-06-18
Added
LITTLE_CODER_EXTRA_EXTENSIONSenv var: layer third-party pi extensions onto the bundled set without forking the installed package (#46). Path-delimited list (:on POSIX,;on Windows —node:path.delimiter) of extension paths. Each entry can be a direct file (e.g. api-ponytail-styleextensions/ponytail.js) or a directory containingindex.ts/index.js(the launcher prefers.ts). A leading~/is expanded; missing paths log a one-line warning to stderr and are skipped (a typo in the env var doesn't kill the session). Survives upgrades — drop the env var into your shell rc once and everylittle-coderrun picks up the extras. Example:LITTLE_CODER_EXTRA_EXTENSIONS=~/.local/lib/node_modules/pi-ponytail/extensions/ponytail.js little-coder. Parsing rules live inbin/extras.mjsso they're unit-testable in isolation (9 cases covering direct-file / dir-index-resolution /index.ts-preference / missing-path warning /~/expansion / multiple entries / whitespace trimming). The launcher-level integration is exercised end-to-end (warning prints for a bad path; valid paths pass through silently to pi as--extension <entry>flags). Closest siblings — third-party skill bundles — are not yet covered;skill-injectstill discovers only<pkgRoot>/skills/tools/*.md, and a follow-up will add the same kind of override.
Notes for upgraders
- No CLI-flag or public-API changes. The new env var is opt-in: unset = identical behavior to v1.9.2. If you were carrying a custom wrapper extension inside the installed npm package (which gets wiped on upgrade), you can drop it and use the env var instead.
[v1.9.2] — 2026-06-18
Fixed
- Width-overflow crash from custom widgets (#48). pi-tui throws
Rendered line N exceeds terminal widthwhenever a custom TUI component emits a line wider than the active terminal — the user saw a 198-char line at width 184 take down the whole session. Root cause was the sub-coder tracker (subagent/tracker.ts): a failed sub-coder'serrorMessageflowed straight into a widget row without any cap, and real-world child-process errors routinely run 150-250 chars (transport error + URL + retry count is enough). The tracker now caps every emitted row to the active terminal width using a new_shared/width.tsutility (visibleWidth+truncateLineToWidth, ANSI-aware so SGR colour codes are preserved through the cut and a final reset prevents bleed).summarizeActivityalso gained a 56-char cap on the failure path (was uncapped) and the running path (was uncapped onpart.name) for defense in depth. The same width-cap is now applied to the plan-mode status panel, the plan-mode indicator, and the branding startup header (which now uses thewidtharg pi passes torender()instead of returning hardcoded-length lines), so a narrow terminal can no longer crash launch either. Newwidth.test.ts(9 cases) covers ASCII / SGR / OSC hyperlink / colour-bleed / the exact issue-48 reproduction shape, andissue-48-repro.test.tsdrives the tracker directly with a 167-char failure at width 184 and asserts no emitted row exceeds the terminal.
Notes for upgraders
- No CLI-flag or public-API changes. If you ever saw
Rendered line N exceeds terminal width (… > …)crash a session — particularly during adispatchcall that errored, or while Plan Mode was orchestrating sub-coders — 1.9.2 fixes it. Third-party pi extensions (e.g.context-mode) that emit their own widgets remain subject to pi's check; if you still see the crash withLoaded pi extensions: <name>listed, the offending widget is in that extension, not little-coder.
[v1.9.1] — 2026-06-08
Fixed
- Plan Mode shortcut moved to
alt+psoshift+tabstays pi's thinking-level cycle (#47). v1.9.0 claimedshift+tabfor Plan Mode by rebinding pi's built-inapp.thinking.cycletoalt+tin~/.pi/agent/keybindings.json. That collided with the muscle memory of every existing pi user —shift+tabis the documented thinking cycle — and pi (≥ 0.79) also surfaced an[Extension issues]warning whenever the rebind hadn't taken yet. Plan Mode now registers onalt+pinstead (unbound by pi, so the extension claims it cleanly with no shadowing), andshift+tabreturns to pi's default behavior. The launcher also performs a one-time cleanup: on first run after upgrade, if~/.pi/agent/keybindings.jsonstill has the v1.9.0 rewrite (app.thinking.cycle: "alt+t"exactly), it is removed; any binding you set yourself is preserved untouched. README and the Plan-Mode indicator ((alt+p to exit)) updated to match.
Notes for upgraders
- No CLI-flag or public-API changes. Plan Mode is now
alt+p(wasshift+tabin v1.9.0).shift+tabis again pi's thinking-level cycle. If you customizedapp.thinking.cycleyourself in~/.pi/agent/keybindings.json, your binding is left alone.
[v1.9.0] — 2026-06-15
Added
- Plan Mode (shift+tab). A Claude-Code-style "research → ask → plan" flow, built as the new
plan-modeextension. Press shift+tab to toggle it (an honey◆ PLAN MODEindicator appears below the input). When it's on, submitting a request does not run a normal coding turn — instead little-coder: (1) decomposes the request into 1-4 exploration tasks, (2) dispatches read-only explorer sub-coders to gather information (their transcripts never enter the main context — only their concise reports survive), (3) generates 1-3 clarifying questions, each with suggested answers plus a free-text "Other" option, asked via the UI, and (4) synthesizes the findings + your answers into a written plan in the chat. Each reasoning phase ("deciding what to explore…", "preparing clarifying questions…") shows an animated spinner with a running m:ss timer. The planning instructions + research are injected into the synthesis turn's system prompt, so the chat shows only your original request and the plan — never the internal scaffolding. A single continuous m:ss timer runs for the whole process (not just the per-sub-coder timers). When the plan is presented, an Approve & implement / Keep planning prompt (arrow keys + enter) gates implementation — only on approval does little-coder start making the changes. Esc (or Ctrl+C) cancels a plan in progress. - Up-arrow prompt history (
prompt-historyextension), persisted across sessions. pi's default editor has no prompt recall; from an empty prompt, ↑ now walks back through your recent prompts (most-recent first) and ↓ walks forward. History is saved to<agentDir>/little-coder-prompt-history.json, so even a brand-new session can recall prompts from earlier runs. Implemented as aCustomEditorsubclass (pi copies its keybindings/autocomplete/submit wiring onto it) usingkeybindings.matchesfor ↑/↓ detection — robust to pi's Kitty keyboard protocol and key-release events — and scoped to recall-from-empty so it never interferes with multi-line cursor movement or the autocomplete dropdown. Edits/writes are blocked during the synthesis turn so plan mode produces a plan, not changes. shift+tab previously cycled the thinking level; pi (≥ 0.79) reserves built-in shortcuts and won't let an extension claim a colliding one, so the launcher rebinds the thinking-level cycle to alt+t in~/.pi/agent/keybindings.json(non-destructively — only when you haven't set your own binding for it), freeing shift+tab for Plan Mode. - Sub-coders (
dispatchtool). little-coder can now spawn isolated child little-coder sessions to research a focused question — single ({ task }) or parallel ({ tasks: [{ label, task }] }, up to 4, concurrency 2 by default, override withLITTLE_CODER_SUBCODER_CONCURRENCY). Children run with the same local-model provider and extensions as the parent (spawned through the launcher headless, not barepi) but are constrained to read + browse-online tools (read, grep, glob, webfetch, websearch, browser, read-only bash) — no edit/write and no recursive dispatch, enforced via the existingtool-gating+permission-gateenv gates. Each child returns a concise report; its full transcript lives in the tool's UI-onlydetailsand never enters the parent model's context, keeping the main window clean. Newsubagentextension (spawn.tsengine, importable by plan mode). - Live sub-coder tracker. A small animated panel above the input shows each running/finished sub-coder with a spinner, status (✓/✗), elapsed time, and current activity (the latest tool call or report snippet), with a diff-guarded ~120 ms repaint. Hidden on non-interactive (benchmark/RPC) runs.
- Session naming + terminal title sync. The session is auto-named from your first prompt (overridable any time with pi's
/name), and the terminal tab title now shows the session name (little-coder · <name>), updating when you switch sessions with/resume. pi's built-in/resumealready lists past sessions for the current directory. - Read-before-edit guard. New
read-guard-editextension: a file must be Read in the current session before it can be Edited — an edit to an unread file is blocked with "File must be read first before edit" and a nudge to Read it (soold_stringmatches exactly). Files you just wrote count as read. Mirrors thewrite-guardenforcement pattern.
Changed
globmatch cap lowered 500 → 100 (extra-tools/glob.ts) to keep results focused for small models. (grepwas already capped at 100.)- Default thinking level is now
mediumfor interactive sessions (pi's default isminimal) — the launcher passes--thinking mediumunless you set a level yourself (--thinking, or a--model …:<level>shorthand) or run headless (--mode/-p). - Auto-named session titles are capped at 4 words, cut on word boundaries (no more mid-word truncation) with a trailing
…when the prompt was longer.
Dependencies
- Bumped bundled pi
@earendil-works/pi-coding-agent0.75.3 → 0.79.4. The "Operation aborted" marker patch (scripts/patch-pi.mjs) still applies cleanly to the new source (verified bypatch-pi.test.mjs). pi 0.79 no longer hoists@earendil-works/pi-tuito the top level, so thedispatchtool's result renderers now build their lines as duck-typed components via the theme (the same patternbrandingalready uses) instead of importing pi-tui primitives — no behavior change.
Notes for upgraders
- No breaking CLI-flag or public-API changes. shift+tab now toggles Plan Mode instead of cycling the thinking level — use alt+t for the thinking-level cycle (the launcher writes this rebinding into
~/.pi/agent/keybindings.json, preserving any binding you've already set). New env varLITTLE_CODER_SUBCODER_CONCURRENCY(default 2) tunes how many sub-coders run at once against your local backend.
[v1.8.4] — 2026-06-08
Added
output-parsernow recognizes LFM2 / Liquid "Pythonic" tool calls (#42). LiquidAI LFM2 models emit tool calls as a Python list wrapped in special tokens —<|tool_call_start|>[Read(path='/a.c'), Bash(command='ls -la')]<|tool_call_end|>— a format neither pi's native path nor the existing fenced/<tool_call>/bare-JSON parsers understood. NewparseLiquidToolCalls()recovers them best-effort: single and double quotes, dict args ({"k":"v"}), list args (['a','b']),True/False/None, ints/floats, commas/parens inside string values, truncated tails (missing)/]/quote), the issue's exact leak shape (start token +[stripped,]<|tool_call_end|><|im_end|>trailing), and the real-world<think>…</think>[calls]shape — all with a precision guard so ordinary prose never trips it. Each recovered call is taggedformat: "liquid"; the extension surfaces a single, accurate diagnostic for that format instead of the futile "use native tool calls" nudge (Pythonic is LFM2's native channel, so nudging would just loop). 20 new parser tests, including one built from verbatim LFM2.5-8B-A1B output.
Fixed / Documentation
- Diagnosed and documented the actual
Failed to parse input at pos N: …<|tool_call_end|>failure (#42). The error is server-side: llama.cpp'schat.cpptool-call parser chokes when the chat template doesn't match it — typically the GGUF's embedded template, which renders tools as a plainList of tools: […]blob without the<|tool_list_start|>/<|tool_call_start|>special tokens the parser expects. Verified end-to-end withLiquidAI/LFM2.5-8B-A1B-Q4_K_M: the embedded template reproduces the error and the tool never runs, while serving with--jinja --chat-template-file LFM2-8B-A1B.jinja(the matching template, with the special tokens) parses calls into nativetool_callsand tools execute normally. New Troubleshooting entry with the exact fix.
Notes for upgraders
- No CLI-flag or public-API changes. If you run an LFM2/Liquid model, serve llama.cpp with
--jinjaand the model's matching chat template (see Troubleshooting). The parser change only adds recovery + a clearer diagnostic for builds that leak the calls as text.
[v1.8.3] — 2026-06-08
Fixed
- User
models.jsonis now found on Windows whenHOMEis unset (#43, thanks @A-M-D-R-3-W). Windows doesn't guaranteeHOME, but it does setUSERPROFILE. The documented fallback~/.config/little-coder/models.jsonwas therefore skipped on Windows and user-defined models never registered.resolveOverridePath()now falls back toUSERPROFILEwhenHOMEis absent (resolution order is unchanged whereHOMEexists:$LITTLE_CODER_MODELS_FILE→$XDG_CONFIG_HOME→$HOME/$USERPROFILE/.config). Path-resolution tests are now platform-neutral viapath.join.
Documentation
- Added an "Any OpenAI-compatible server (e.g. MLX / omlx)" section to the model-configuration docs (#40). little-coder registers providers from
models.jsonrather than from pi's standalone picker extensions, so an omlx/MLX server is added by declaring a provider entry (any OpenAI-compatible/v1endpoint works the same way), not by installing its pi picker. The README now shows the exact~/.config/little-coder/models.jsonblock.
[v1.8.2] — 2026-05-25
Fixed
- Minimal user
models.jsonentries no longer crash startup withCannot read properties of undefined (reading 'input')(#36). The shippedmodels.jsondeclares every field —id,name,reasoning,input,contextWindow,maxTokens,cost— but a user override that omitted e.g.name/maxTokens/costwas passed through unchanged to pi's registry, which then exploded deep inapplyModelOverridewhen it tried to readmodel.cost.input.llama-cpp-providernow fills in the same defaults pi uses for built-in models (name = id,reasoning = false,input = ["text"],contextWindow = 32768,maxTokens = 4096, zero-cost) so a minimal entry — justidplus the provider'sbaseUrl/apiKey— works. User-supplied values still win over defaults; unknown extra fields (e.g._launch) are preserved. A model entry that omitsidis now flagged with a precise error in the source diagnostics instead of crashing pi. NewfillModelDefaultshelper, plus regression tests using the exact entry shape from the issue report. temperature' is not supported with this modelagainst Copilot GPT-5.x / OpenAI o-series (#33).benchmark-profileswas injectingtemperature: 0.3fromdefault_model_profileinto every outgoing chat-completions payload, but hosted reasoning models hard-reject the parameter with a 400. The temperature injection is now gated on the provider: it ships on forllamacpp,ollama, andlmstudio(the providers it was tuned for) and is skipped for everything else. New env varLITTLE_CODER_TEMPERATURE_PROVIDERS=foo,barreplaces the default list when you bring your own local provider (e.g.vllm). New exported, testedproviderAcceptsTemperature(); end-to-end test firesbefore_agent_start+before_provider_requestand asserts the copilot path returns no payload mutation.
Notes for upgraders
- No CLI-flag or public-API changes. If you previously relied on temperature 0.3 reaching a non-local provider via the default profile (uncommon — most hosted providers reject it), add that provider name to
LITTLE_CODER_TEMPERATURE_PROVIDERS.
[v1.8.1] — 2026-05-23
Fixed
globno longer exhausts memory on a recursive search from a huge root. The tool capped matches at 500 but never bounded the walk: run from a home directory (or any tree with macOSLibrary, caches, ornode_modules),fs.globrecursively descended everything and its internal traversal state grew until the Node process ran out of heap — a host-memory crash (Ineffective mark-compacts near heap limit), entirely distinct from the model's context window (the read-guard / window machinery operates on tool results in tokens; this died mid-walk, before any result existed). The walk is now bounded two ways: heavy/irrelevant directories (node_modules,.git,dist,.cache,Library,venv,target, …) are pruned — never descended — and a hard scan budget (200 000 entries) halts the walk through the one hookfs.globcalls per entry (exclude), since it exposes no signal/abort. When results are cut short the output says so, so the model narrows its search. New unit-testedglobFiles/renderGlobOutcomehelpers (.pi/extensions/extra-tools/glob.ts), verified to prunenode_modules(0 descent) and to halt at the scan budget.
Notes for upgraders
- For a focused search, pass a
path(a project subdirectory) instead of globbing from a home directory. Hidden directories continue to be skipped byfs.globas before.
[v1.8.0] — 2026-05-23
little-coder now auto-detects the llama.cpp server's live context window at startup and registers the model with it, so a llama-server -c 131072 shows 128k instead of the declared default — no config edit. This completes v1.7.0: the budget already followed the registered window; now the registered window itself comes from the running server.
Added
- Live context-window detection for llama.cpp. On startup
llama-cpp-providerGETs the server's/propsendpoint, reads its actualn_ctx, and registers the model with that window in place of the staticcontextWindowinmodels.json. The TUI context readout, read-guard's overflow trim, and the skill/knowledge budgets all then track the server's real window — bumpllama-server -cand little-coder follows, nomodels.jsonor settings edit. The/propsURL is derived from the provider baseUrl by stripping/v1(llama-server serves it at the root); the value is read fromdefault_generation_settings.n_ctx. New tested helperspropsUrlFor/contextWindowFromProps/probeContextWindow, validated end-to-end against a live-c 131072server (→ 131072).- Best-effort and safe: 1.5 s timeout,
llamacppprovider only, and ANY failure (server down, no/props, non-JSON, timeout) silently falls back to the declared window — startup is never blocked or broken. - Env knobs:
LITTLE_CODER_NO_CTX_PROBE=1disables the probe (offline / CI);LITTLE_CODER_LLAMACPP_PROPS_URLoverrides the/propsURL for non-standard setups;LITTLE_CODER_CTX_PROBE_TIMEOUT_MStunes the timeout.
- Best-effort and safe: 1.5 s timeout,
Notes for upgraders
- This adds one best-effort HTTP GET to the llama.cpp
/propsendpoint at launch (only for thellamacppprovider). If your server/proxy doesn't expose/props, behaviour is unchanged — the declaredmodels.jsoncontextWindow(default 32768) is used. SetLITTLE_CODER_NO_CTX_PROBE=1to skip the probe entirely. - No CLI-flag or public-API changes.