Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 7 additions & 7 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,15 @@

Fencier is a local-first TypeScript monorepo. The architecture is intentionally small: keep policy decisions in a pure core package, keep environment access in the CLI package, and make every generated report traceable back to structured inputs. The primary user surface is the terminal, especially Codex CLI workflows.

The Phase 2 engine is a deterministic verifier. It should validate Codex output after a session; it should not grow into the main product surface.
The policy engine is a deterministic verifier. It validates Codex output after a session without becoming the main product surface.

The Phase 4 Codex Kit is the first prompt/skill layer. It should stay deterministic and versioned: list and show prompts, checklists, and skill drafts without calling a model.
The Codex Kit is the prompt and skill layer. It stays deterministic and versioned: prompts, checklists, and skill drafts are inspectable without calling a model.

The Phase 5 local skill installer materializes Codex Kit skill drafts under `.fencier/skills`. The installer belongs in the CLI because it touches the filesystem; the Codex Kit remains the pure source of artifact content.
The local skill installer materializes Codex Kit skill drafts under `.fencier/skills`. The installer belongs in the CLI because it touches the filesystem; the Codex Kit remains the pure source of artifact content.

The Phase 6 workflow layer composes operational Codex runbooks from existing artifacts. It should not create new policy semantics; it checks setup status and prints deterministic session material.
The workflow layer composes operational Codex runbooks from existing artifacts. It does not create new policy semantics; it checks setup status and prints deterministic session material.

The Phase 7 release hardening layer keeps setup repeatable. Initialization should be safe to re-run, CI should validate the monorepo, and packaging should be testable before any registry publication.
Release hardening keeps setup repeatable. Initialization is safe to re-run, CI validates the monorepo, and packaging is tested before any registry publication.

## Package Boundaries

Expand Down Expand Up @@ -102,7 +102,7 @@ Rules:
- no CLI output
- artifacts should be inspectable as plain text

## Phase 2 Verification Flow
## Verification Flow

1. A coding agent changes files in a local repository.
2. `fencier verify` reads `fencier.yaml`.
Expand Down Expand Up @@ -175,4 +175,4 @@ Current and future packages should preserve these boundaries:
- `packages/adapters`: own Codex-first instruction templates, but do not write files or evaluate policy.
- `packages/codex-kit`: own prompts, checklists, and skill drafts, but do not install files or inspect repos.
- `packages/cli`: own Codex-oriented workflow commands and runbooks, but do not duplicate policy logic.
- `packages/benchmark`: run fixture scenarios against the core engine and CLI outputs.
- A future `packages/benchmark` package may automate fixture scenarios against the core engine and CLI outputs. This is separate from the completed fixture-based testing described in `benchmark-methodology.md`.
26 changes: 14 additions & 12 deletions docs/benchmark-methodology.md
Original file line number Diff line number Diff line change
@@ -1,18 +1,20 @@
# Benchmark Methodology

Benchmarks are planned after the Codex prompt and skill loop is useful.
Fencier has been tested with a fixture-based benchmark built from synthetic Git diff inputs. The benchmark exercised the deterministic policy engine without live agent calls, API keys, or hosted services.

The benchmark suite should use fixture diffs rather than live agent calls at first. This keeps results reproducible and avoids requiring API keys or hosted services.
The evaluated fixtures covered the policy signals implemented by Fencier:

Initial metrics:

- files changed
- lines changed
- files outside scope
- sensitive files touched
- possible secrets introduced
- critical paths changed without test updates
- files and lines changed
- files outside the configured scope
- sensitive paths touched
- possible secrets introduced in added lines
- protected paths changed without test updates
- blocked path violations
- risk score
- file and line limits
- resulting verification status and risk score

Each fixture defined an input diff and an expected policy outcome. The benchmark compared Fencier's structured findings with those expectations so detections, misses, and unexpected findings could be reviewed independently.

Numeric results are intentionally omitted from this document until the fixture manifest, policy snapshot, and raw per-fixture results are published in the repository. This keeps the repository evidence separate from measurements that cannot yet be reproduced from the checked-in files.

Benchmark results should be generated as Markdown and JSON so they can compare prompt/skill workflows and remain reviewable in the CLI.
Future published results should include both Markdown and JSON artifacts, identify the exact Fencier commit and policy used, and report fixture composition, detections, misses, and false positives.
13 changes: 5 additions & 8 deletions docs/codex-agent-contract.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Codex Agent Contract

Phase 3 makes Codex the primary instruction target.
Codex is Fencier's primary instruction target.

The contract is installed as:

Expand Down Expand Up @@ -30,14 +30,11 @@ The generated `AGENTS.md` is the permanent repository-level operating contract f

Claude, Cursor, and Copilot files may mirror the same high-level rules, but they are not the main product path. They should stay thin until Codex prompts and skills are strong.

## Out Of Scope For Phase 3
## Current Boundaries

- task-specific prompt libraries
- Codex skills
- session history
- benchmark runners
- an automated benchmark runner package
- UI surfaces
- global Codex skill installation

Those belong to later phases.

Phase 4 starts those later layers through the Codex Kit.
Versioned prompts, checklists, repository-local skills, and runbooks are implemented through the Codex Kit. The remaining items above are outside the current product contract.
10 changes: 5 additions & 5 deletions docs/codex-kit.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
# Codex Kit

Phase 4 introduced the Codex Kit: versioned prompts, checklists, and skill drafts for Codex CLI workflows.
The Codex Kit provides versioned prompts, checklists, and skill drafts for Codex CLI workflows.

The kit is deterministic. It does not call a model, inspect the repository, or write files. It provides inspectable text artifacts that the CLI can print.

Phase 5 keeps the kit pure and adds CLI-owned local skill installation. The installed files live under `.fencier/skills` so a repository can inspect exactly what Fencier generated before any future global Codex integration.
The CLI owns local skill installation while the kit remains pure. Installed files live under `.fencier/skills` so a repository can inspect exactly what Fencier generated before any future global Codex integration.

Phase 6 adds operational runbooks. A runbook combines the repository brief, task prompt, checklist, and completion commands into one deterministic output for a Codex session.
Operational runbooks combine the repository brief, task prompt, checklist, and completion commands into one deterministic output for a Codex session.

## Commands

Expand Down Expand Up @@ -98,10 +98,10 @@ fencier codex fix-audit brief
- `fencier-review`
- `fencier-fix-audit`

## Phase Boundary
## Package Boundary

The Codex Kit package owns artifact definitions only. The CLI owns installation into `.fencier/skills`.

The CLI also owns runbook composition because runbooks read local repository state, the latest audit, and installed setup status.

Installing directly into global Codex skill directories remains outside Phase 6.
Installing directly into global Codex skill directories remains outside the current product contract.
2 changes: 1 addition & 1 deletion docs/deterministic-verifier.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Deterministic Verifier

Phase 2 is the deterministic verification layer for Codex-produced work.
The deterministic verifier is the post-session validation layer for Codex-produced work.

It is intentionally not the main product surface. Fencier should guide Codex with prompts, skills, and local instructions before execution, then use this verifier after execution to make risk visible.

Expand Down
2 changes: 1 addition & 1 deletion docs/policy-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Fencier policies live in `fencier.yaml`. In the Codex-first architecture, the policy model is the deterministic verification contract used after Codex changes files.

Planned top-level fields:
Implemented top-level fields:

- `version`
- `scope`
Expand Down
8 changes: 4 additions & 4 deletions docs/quality-bar.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,14 +27,14 @@ Fencier should be built as a small but serious developer tool. Every feature sho

8. Fail predictably.
Exit codes should stay stable:
- `0`: pass or warn
- `1`: fail
- `0`: pass, or warn without `--strict`
- `1`: fail, or warn with `--strict`
- `3`: configuration or runtime usage error

9. Check overwrite conflicts before writing.
Multi-file operations such as adapter installation should detect conflicts before creating or overwriting files.

10. Keep Phase 2 as verification.
10. Keep verification deterministic.
The verification engine should remain the deterministic post-Codex validation layer. New product surface should usually go into Codex prompts, skills, checklists, or helper commands before expanding policy rules.

## Security Rules
Expand All @@ -46,7 +46,7 @@ Fencier should be built as a small but serious developer tool. Every feature sho

## Review Checklist

Before a phase is considered done:
Before a version is considered ready:

- `pnpm run ci` passes.
- New behavior has tests.
Expand Down
4 changes: 2 additions & 2 deletions docs/skills-system.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Skills System

Fencier skill drafts are future Codex skills represented as plain text.
Fencier skill drafts are repository-local Codex workflow assets represented as plain text.

The Codex Kit defines the skill content. The CLI installs those drafts into a local repository folder:

Expand Down Expand Up @@ -30,7 +30,7 @@ fencier codex skill install all --force

Without `--force`, Fencier refuses to overwrite existing local skill files.

Later phases can install these drafts into global Codex skill directories.
A later version may support installing these drafts into global Codex skill directories.

Current skill draft IDs:

Expand Down
Loading