AI coding agent built from scratch in pure Go — CellGrid TUI, native MCP client, web scraping, and LSP diagnostics in a single binary.
Custom CellGrid frame-buffer renders markdown directly in the terminal — no external renderer dependency.
- Bold, italic,
inline code, headings, blockquotes, ordered/unordered lists (nested arbitrary depth) - Fenced code blocks with
\t→ spaces conversion, indentation preservation - Tables with smart column-width alignment, box-drawing characters (
│ ─ ┤ ├ ┼) - Horizontal rules, links, mixed content
- Incremental streaming:
parseMarkdown()on every render tick — no sudden jump from raw markdown to styled output - 10+ roundtrip tests verify text content survives parse → render → extract
- Auto-save conversations on exit (
~/.tinycode/sessions/TUI-*.json), with title/preview/metadata - Resume with
--resume=TUI-20260607-235959 - List sessions with
--list-sessions(shows title, message count, last active time) - AI-generated session titles: title hidden agent auto-names conversations via LLM. Falls back to first user message. (2deee5e)
- Input history: Up/Down arrows to recall previous inputs, Esc to exit browse mode
- Status bar shows
hist 3/5during history browsing
- Incremental CellGrid Rendering — msgDirty/msgRowCount tracking. View() skips unchanged messages, only re-renders from first dirty onward. CellGrid no longer Reset() every frame. Benchmark: ~2.3ms regardless of message count (10/50/100 tested). (ffbbf22)
- Reasoning folding: click
[+]/[-]markers to expand/collapse LLM reasoning blocks. Multi-step agent loop now creates one assistant message per step, each with its own[-]reasoning +→ Calling tools:list. (c5f854f) - Tool call display:
→ Calling tools:with bullet list during tool execution. Duplicate tool calls in the same step no longer appear twice (agent.go fired OnToolCall twice per tool — fixed in cc5a224) - Character-level selection: drag-select any range of viewport text, Ctrl+C copies rendered text (not raw markdown). Range check prevents single-click selects from triggering copy. (ea6711c)
- [Copy] buttons: click to copy assistant response to clipboard. Fixed: button row now counted in msgRowCount so subsequent messages don't overlap. (f9143cc). Fixed: button stops working after incremental render — activeButtons preserved per MsgIdx across renders. (8dffd64)
- Spinner:
⣾ ⣽ ⣻ ⢿ ⡿ ⣟braille spinner in status bar during streaming. Tick pipeline kept alive even during idle. Spinner continues across intermediate steps. (42b8794, 9a0f55b, afc14e1) - Auto-scroll: viewport follows streaming output, pauses when user scrolls up
- Status bar: mode icon, model name, provider, token/tool/msg counts, session duration, transient status messages
- Frame verification — the rendered frame is a test artifact, not only an assertion target: 27 plain-text golden frames plus one raw ANSI frame under
tui/testdata/golden/(regenerate withgo test ./tui -run Golden -update), andmake test-tui-visual(TUI_SHOT=1), which renders the scenarios to PNGs, starts the built binary on a real 80x24 PTY (and once on a size-less one) and replays that live stream into a screen buffer that is rendered too. The mechanism is the nested tuiprobe module (pinned ingo.mod, released astuiprobe/v0.1.0); the fixtures and the judgments stay here. Reference: docs/tui-verification.md.
- ReAct loop with tool calling support (24 tools: bash, read_file, write_file, search_files, edit, apply_patch, git_*, web, LSP, task, task_collect, todo, sandbox_allow, load_skill, skill_manage)
- 6 agents: plan (read-only whitelist), build (full access), explore (read_file + search_files only), general (full execution sub-agent), compact (history compression), title (session naming)
- Permissions engine:
Rulesetwith last-match-wins{action, resource, effect}rules replacing DeniedTools/AllowedTools. Supports whitelist (*: deny+ specific allows) and blacklist (*: allow+ specific denies). - Task tool: Delegate to sub-agents via
task({agent, goal}). Sync mode (block until done) or bg mode (returns task_id, collect withtask_collect). Sub-agent steps don't count against parent's step budget. - OnStepDone callback: after each step's tools complete, the agent fires
OnStepDone— the TUI creates a new assistant message for the next step. Each step gets its own reasoning + tool calls display. - Agent integration test framework: 13 tests using MockLLM step-by-step
- Streaming reasoning + text deltas
- Tool call lifecycle displayed in real-time
- Multi-turn history compression: Hermes-style head/tail/middle summarization
- 1M token context (DeepSeek V4 Flash) with automatic threshold lowering on
context_length_exceedederrors
- Config:
"lsp": { "enabled": true }in config.json (default: disabled). 7 supported languages with auto-detection: Go (gopls), Python (pyright), TypeScript/JS (typescript-language-server), Rust (rust-analyzer), C++ (clangd), Java (Eclipse JDT). - 4 tools exposed to LLM:
lsp_definition(→Definition at path:line:col),lsp_references(→Found N references),lsp_hover(→ type info + docs),lsp_symbols(→ all symbols in file). All requirefile_path,line,character(0-indexed). - Architecture:
lsp.Init(<project root>)records the workspace (the language server is rooted at the project, not at the session directory) and the server starts lazily on the firstSyncFile. One reader goroutine owns the stream and dispatches responses to per-id channels pluspublishDiagnosticsto a diagnostics channel; the connection is reused by every later tool call and shut down onClose(). If no persistent client is available, a tool falls back to a one-shotexec.Commandserver for that call. - Incremental Diagnostics —
SnapshotBaseline(path, content)captures diagnostic state from the pre-edit bytes before edit/write_file/apply_patch,GetNewDiagnostics(path, content)computes the delta from the bytes just written. Tools report only new errors via LSP — LLM sees focused feedback. (a2e3e07) - TUI Error Tracking — an in-memory registry (severity-1 diagnostics per file) is refreshed off the event loop after each tool result and on the spinner tick, so the status bar
errors: Nand/diagnosticsreflect the live state. (290818a) - Mock test framework —
io.Pipebased, no LSP server required; covers all 4 tool types, concurrent request correlation and reader-death handling. (8065ae5) - Integration tests —
make test-lspruns the suite against a realgopls(the Nix devShell provides it,make install-goplsinstalls the pinned version elsewhere); CI runs them in a dedicated job. They cover a valid file reporting no errors, a broken file reporting its undefined symbol, and the error clearing after a fix. - Limitation: the first call of a session pays the server startup cost (~500 ms); subsequent calls reuse the connection. The one-shot fallback (no persistent client) still pays it per call.
- TodoStore: In-memory task list with CRUD (create/read/update/merge/delete/summary). Enforces one
in_progress, max 256 items, max 4000 chars per task. - todo tool: OpenAI function-calling schema, registered in all modes. LLM calls it to plan and track multi-step work.
- Compression protection: After context compression, active todo items (pending + in_progress) are re-injected so the LLM doesn't redo completed work.
- Housekeeping mute: When all tool calls are
todo, the model's text reply is suppressed — no noise, just progress markers. - Session recovery: On
--resume, reverse-scans history for the latest todo result and restores the store. - TUI rendering:
▾ Todo (2/6)with[x]completed,[>]in_progress,[ ]pending,[~]cancelled markers.
- edit tool: Search/replace editing (old_string + new_string). 7 fuzzy strategies (exact → line-trimmed → ws-normalized → indent-flexible → escape-normalized → unicode-normalized → block-anchor) + indentation correction. Validates uniqueness. Multiple edits per call. LSP integration. 14 tests. (d067156, 34d2c17)
- apply_patch tool: V4A multi-file patch format. Supports UPDATE (line-level -/+ hunks), ADD (create files), DELETE (remove files). Two-phase execution: validate all, then apply. Multi-file in one call. 9 tests. (9045176)
- write_file preserved for creating new files and full rewrites. Three tools form a complementary editing system.
- Native MCP client: Connect to MCP servers via stdio (subprocess) or HTTP (remote endpoint). Config-driven via
mcp_serversin config.json. - Auto-discovery: On startup, connects to all servers, calls
tools/list, and registers each discovered tool as an independentmcp_<server>_<tool>agent tool with its original JSON Schema. - Transport: stdio (exec.Command with pipes, Content-Length framing) or HTTP POST (configurable headers, JSON-RPC 2.0).
- Resources:
resources/listandresources/readsupport for MCP resources. - Security: SSRF protection for HTTP transport — blocks private IPs (RFC 1918/loopback/link-local), fails closed on DNS failure. Localhost allowed for dev use.
- Graceful degradation: Server connection failure logs a warning and skips that server — other servers still work. 22 tests. 359 total. (e31c08b, cc1ba8e, 7b4fe75)
- web_search: Searches the web using DuckDuckGo Lite (zero config, no API key). Optional SearXNG fallback configured via
searxng_urlin config.json. Returns numbered results with title, URL, description. (5c90f86, 4cb7ce1, 405bc87) - web_extract: Fetches and extracts web page content as Markdown. 5-level fallback chain: direct HTTP → Cloudflare bypass (UA retry) → Google Cache → Wayback Machine (CDX + id_ format) → Chromium headless. SSRF protection: DNS resolution + IP blacklist (RFC 1918, loopback, cloud metadata). LLM summarization: content >5000 chars auto-summarized via provider. The browser is discovered once and probed with
--version, preferringCHROME_PATH/CHROME, then the system browsers, then the Playwright cache, so a broken launcher is skipped rather than handed to the extractor. The headless run uses a throwaway browser profile, so extraction never touches the profile your own Chrome uses. (5c90f86, 4cb7ce1, 4bf27e5)
- GitHub Actions: Two workflows — main.yml (build + lint + test on push/PR) and release.yml (cross-compile + GitHub Releases on tags v*)
- main.yml jobs:
ci(build, gofmt gate, vet, tests,-race, repeated run),lsp(real gopls),browser(real Chromium),tui-visual(frame PNGs + the built binary on a PTY),cross(linux/amd64, linux/arm64, darwin/arm64 build + vet) andstaticcheck - Makefile improvements: test target preserves exit code with pass/fail message; releases target cross-compiles all platforms + .tar.gz archives
- 710 test functions + 10 fuzz targets across all packages, counted with
grep -rn '^func Test' --include=*_test.go . | wc -landgrep -rn '^func Fuzz' --include=*_test.go . | wc -l(the command is part of the record, and so is the badge below:CODEBASE.md→ Testing lists every count with its measurement) - Annotations: the jobs carry one
ubuntu-latestmigration notice each, plus thesetup-chrome@v1Node 20 warning onbrowserandtui-visual; the gate is no new annotations, not zero
Every gated target skips itself without its environment variable, so a plain make test
never launches a browser, a language server or a terminal device. Run them deliberately:
| Check | Command | Needs |
|---|---|---|
Linux-only cases (the openat2 containment tests, the cgroup reaping) |
docker run --rm -v "$PWD":/w -w /w golang:1.27 go test ./tool/ |
a running Docker daemon; the kernel inside needs ≥ 5.6 for openat2 |
| Real language servers | make install-gopls, make install-tsls (TypeScript 5, optional), then make test-lsp (LSP_TEST=1) |
a Go toolchain; npm for the TypeScript case, which skips with the reason without it |
| Real browser smoke tests | make test-browser (BROWSER_TEST=1) |
a Chromium/Chrome: a system install, CHROME_PATH=…, or npx playwright install chromium |
| TUI frames and the PTY smoke test | make test-tui-visual (TUI_SHOT=1) |
/dev/ptmx and a fresh bin/tinycode (Chromium only for the browser renderer) |
Two failures that are the environment, not the code:
- Chromium may not start in a restricted sandbox (the agent sandbox denies the user
namespace, the profile directory or
/dev/ptmx). Run the browser and PTY targets on a normal host, or leave them to thebrowserandtui-visualCI jobs. --dump-domcan hang with a full desktop Chromium — measured on macOS with the Playwright "Google Chrome for Testing" build:--dump-domproduced nothing in 120 s, whilechrome-headless-shellfrom the same revision dumped the page in about a second (issue #43). The extractor's--dump-dompath therefore prefers the Playwright headless shell and falls back to the full browser, while therodpath keeps preferring the full browser;CHROME_PATH/CHROMEstill win over both, so a workflow that pins its browser (CI'ssetup-chrome) is unaffected.
- SKILL.md-based discovery — three-layer scan: embedded (skill/builtin/) → ~/.tinycode/skills/ → project .tinycode/skills/ (upward search). Later sources override earlier. (cbd6db3)
- /skill command in TUI — /skill lists available skills; /skill loads full SKILL.md content as system message. (cbd6db3)
- Skill index auto-injected into system prompt at startup. Startup shows "13 tools, 2 skills loaded". (8fa8800)
- 2 builtin skills: code-review, git-commit (as markdown files in skill/builtin/)
- 6 agents: plan, build (primary) + explore, general (subagents) + compact, title (hidden)
- 11 new tests across skill package and tui package — 359 tests total. (cbd6db3)
┌──────────────────────────────────────────────────────────┐
│ TinyCode │
├──────────────────────┬───────────────────┬────────────────┤
│ TUI (Bubble Tea) │ Agent Layer │ Tool Layer │
│ │ (ReAct Loop) │ (24 tools + MCP)│
│ CellGrid │ │ │
│ Viewport │ Plan (primary) │ bash │
│ Input Area │ Build (primary) │ read_file │
│ Status Bar │ Explore (sub) │ write_file │
│ Command Palette │ General (sub) │ edit │
│ Todo Display │ Compact (hidden) │ apply_patch │
│ Reasoning Fold │ Title (hidden) │ search_files │
│ │ │ task │
│ │ Registry: │ todo │
│ │ Get/Set/Switch │ memory │
│ │ ToolAllowedFor │ load_skill │
│ │ Subagent→task │ skill_manage │
│ │ │ lsp_* (4) │
│ │ │ sandbox_allow │
│ │ │ web_search │
│ │ │ web_extract │
└──────────────────────┴───────────────────┴────────────────┘
6 agents configured in agent/config.go, managed by agent/registry.go:
| Agent | Mode | Hidden | Tools | Steps | Purpose |
|---|---|---|---|---|---|
| plan | primary | read_file, search_files, git_, web_, lsp_*, todo, load_skill | 20 | Read-only analysis (whitelist, no bash) | |
| build | primary | * (all 24 tools) | 50 | Full access implementation | |
| explore | subagent | read_file, search_files | 15 | Fast read-only code search | |
| general | subagent | * except {task, task_collect, skill_manage} | 20 | Full-execution parallel sub-agent | |
| compact | primary | ✅ | (no tools) | 1 | History compression |
| title | primary | ✅ | (no tools) | 1 | Session title gen |
- Primary agents: user-switchable via Tab or /plan /build
- Subagents: invoked via
tasktool with independent ReAct context - Hidden agents: pure LLM calls (no tools), used internally
Component → []CellChunk → wordWrap → Grid.AppendChunk → Grid.Render() → viewport
CellGrid— flat array ofCell{ Rune, Style, Width }, auto-grows as content is addedCellChunk— struct withText stringandStyle CellStyle(Bold, Italic, Underline, Fg, Bg)wordWrap— splits text at width, preserves leading spaces, returns[]CellChunkFill— appliesSelectionStyleto a rectangular cell rangeExtractText— returns plain text within a range, handles CJK multi-cell characters- Incremental rendering: msgDirty/msgRowCount tracking. First dirty → end. ~2.3ms constant.
type Tool struct {
Name string
Description string
Parameters map[string]any
Execute func(ctx, args) (string, error)
}Line-level editing (3 tools):
write_file— create new files / full rewritesedit— search/replace with 7 fuzzy strategies + indentation correctionapply_patch— V4A multi-file patch (UPDATE/ADD/DELETE)
Sandbox:
File effects are restricted by a policy frozen onto each run — one mode
(read-only / workspace-write / danger-full-access) and one list of writable
roots, read by the path fence, the command boundary and the plan-mode guard
alike, so no two of them can disagree about what is writable.
- Path fence — the file tools resolve and compare paths against the roots, and
where the kernel can, it re-checks the decision at the point of use
(
openat2 RESOLVE_BENEATH; theO_NOFOLLOWcomponent walk on macOS). - Command boundary — on by default where the host can apply it (Landlock on
Linux),
bashruns under the kernel boundary: a write outside the roots is refused. On a host with no mechanism the default is off and/sandboxsays so; forcing it on there refuses the command rather than running it unconfined. The platform user cache root is granted to both the fence and the command so a confined toolchain can writeGOCACHEand friends. - Approval dialog — Allow once (one call, nothing cached), Allow
session (the session file, restored on resume), Always allow (the user
config, named in the dialog and revocable with
--revoke-grant), Deny. It auto-shows when a tool is refused, and aread_filecan trigger it too. /sandboxreports the containment level, the command-boundary switch and the effective roots; every refusal is one marker with one shape, carrying what kind of recovery is possible.
Design, the launch protocol, grant lifetimes, the refusal vocabulary and what is
deliberately not promised: docs/sandbox.md. Command
confinement defaults on where the kernel can enforce it (Linux with Landlock) and
off where it cannot; the per-platform default is stated in /sandbox.
Permissions: ToolAllowedFor(cfg, toolName) — checked before every tool execution. Plan mode denies write/git/task/skill_manage.
type LLMProvider interface {
Chat(ctx, ChatRequest) (*ChatResponse, error)
Name() string
}- DeepSeek (default): streaming SSE support,
deepseek-v4-flash - MockLLM: step-by-step scripted responses for agent loop testing
- ProviderRegistry: switch providers at runtime via Tab
History threshold: 50% of context window
Head: system + first 2 exchanges (preserved)
Tail: last 2 exchanges (preserved, anchored on latest user msg)
Middle: → LLM summarization → [COMPRESSED HISTORY] system message
Active TODO injected: [ACTIVE TODO ITEMS] after compression
/compresscommand for manual trigger: the summarizer runs off the event loop (the UI stays responsive), refuses while a run is active, and Ctrl+C cancels a stalled request- The summarizer call runs under the caller's context — a cancel or a provider failure leaves the history untouched and is reported in the status bar instead of silently reporting "nothing to do"
- Auto-recovery:
context_length_exceeded→ lowers threshold - Todo protection: active items re-injected in compressed output
User Input (textarea)
↓
ChatMsg → agent.Run() → ReAct Loop
│ ├── LLM provider (streaming SSE)
│ ├── Tool execution (permissions checked)
│ └── No tool call → return final answer
↓
streamCh (buffered 200)
↓
TUI Update() → ToolCallMsg / StreamMsg / StreamDone
↓
TUI View() → Component.Render() → CellChunks → CellGrid
↓
viewport.SetContent() → terminal display
agent/ Agent loop, LLM provider, context compression, registry
config/ Config loading (JSON, env, CLI flags)
docs/ Reference documents (sandbox design, TUI visual harness)
internal/netsafe/ Shared SSRF policy (blocked IPs, pinned-IP client, redirect checks)
lsp/ LSP client (gopls), diagnostics, Formatter, touch
session/ Session persistence (JSON files, metadata, listing, fork)
skill/ SKILL.md discovery (3-layer), Load/LoadOnce/CRUD
tool/ Tool definitions (24 tools + MCP: edit, todo, skill, LSP, web, mcp)
tui/ Bubble Tea TUI (CellGrid, components, key/mouse, cmd palette)
types/ Shared types (Message, ToolCall, StreamCallbacks)
main.go CLI entry point with cobra
| Package | Purpose |
|---|---|
bubbletea |
TUI framework, event loop |
bubbles/viewport |
Viewport widget |
bubbles/textarea |
Input textarea |
bubbles/spinner |
Loading spinner |
lipgloss |
ANSI style management |
go-runewidth |
CJK character width calculation |
go-openai |
LLM provider (OpenAI-compatible) |
goldmark |
Markdown parser |
cobra |
CLI flag handling |
A reproducible development environment is provided via a Nix flake
(flake.nix) plus a direnv integration (.envrc). The toolchain is
pinned so make build / make test / make lint behave identically
across machines.
Toolchain summary
| Tool | Version | Notes |
|---|---|---|
| Go | 1.27 (module go directive) |
Pinned to the 1.27 line by the Flake (pkgs.go_1_27); CI uses Go 1.27 |
| Make | — | build, run, test, lint, cross-compile |
staticcheck |
optional | Used by make lint (best-effort, ` |
Activate the environment
# With direnv (recommended) — run once, then cd into the repo:
direnv allow
# Or without direnv:
nix developThe dev shell provides the Go toolchain, gopls (LSP), gofumpt
(formatter), and git. Once active, the usual commands work directly:
make build # static binary at ./bin/tinycode (CGO_ENABLED=0)
make test
make lint
make test-tui-visual # TUI screenshots + PTY smoke (TUI_SHOT=1)The Makefile remains the canonical build/release path; the flake pins the toolchain so it behaves identically everywhere.
The TUI visual layers (golden frames, PNG screenshots, the PTY smoke tests and the live-stream replay) have their own reference: docs/tui-verification.md.
The sandbox — modes, the frozen per-run policy, the path fence, the command boundary and its launch protocol, grant lifetimes, the refusal vocabulary, and what is deliberately not promised — has its own reference: docs/sandbox.md.
What TinyCode does not have yet, verified entry by entry against the source and recorded once so it does not have to be re-derived, is inventoried in docs/roadmap.md. It is a record rather than a commitment: an entry becomes work when it is picked up, tracked in issue #168.
The feature log, the current hardening work and the commit hashes they landed with live in CHANGELOG.md. All planned features of the original roadmap are implemented; the MCP client supersedes the old plugin-system proposal.
Built with ❤️ and Go