diff --git a/.github/skills/skill-creator/SKILL.md b/.github/skills/skill-creator/SKILL.md new file mode 100644 index 0000000..eb6370c --- /dev/null +++ b/.github/skills/skill-creator/SKILL.md @@ -0,0 +1,153 @@ +--- +name: skill-creator +description: > + Create new agent skills, improve existing ones, and evaluate their performance with test cases. + Use when a user wants to turn a workflow into a reusable skill, refine a SKILL.md, optimize + triggering and output contracts, or benchmark a skill against a baseline. +--- + +# Skill Creator + +A meta-skill for turning workflows into durable, testable agent skills and for iterating on them until they are reliable. + +This is a distilled synthesis of the most authoritative patterns in the agent-skills ecosystem, especially the Anthropic `skill-creator` guidance and the broader agent-skills spec used across the community. + +## When to Use + +- User wants to "turn this workflow into a skill" or "create a skill for X" +- Existing `SKILL.md` needs editing, restructuring, or quality improvement +- The skill description needs better triggering behavior +- The task benefits from evals, benchmark comparisons, or iterative tuning +- User wants to formalize a repeatable process into a reusable capability + +## Process + +### 1. Capture intent + +Start by defining the skill's job in one sentence: what it should enable the agent to do, when it should trigger, and what the expected output is. + +Ask clarifying questions if needed about: +- the core workflow to turn into a skill +- likely user prompts that should trigger it +- required inputs and expected output format +- edge cases, dependencies, and scope boundaries + +### 2. Research and interview + +Before writing the final skill, gather enough context to make it general and robust: +- inspect similar skills or patterns in relevant repos +- check the exact workflow and any existing examples +- identify common failure modes and edge cases +- decide whether this should be a deterministic workflow, a research workflow, or a subjective creative workflow + +### 3. Draft the SKILL.md + +Write a clear, usable skill file with: +- valid YAML frontmatter with `name` and `description` +- a concise but specific description that includes trigger contexts +- `## When to Use` bullet points with concrete situations +- an actionable `## Process` section with ordered steps +- explicit output format or example results +- explicit boundaries and scope limits + +Prefer imperative language and focused responsibilities. + +### 4. Structure for progressive disclosure + +Keep the skill readable and token-efficient: +- metadata is the short description used for triggering +- the `SKILL.md` body contains the process and patterns +- supporting references or bundled scripts live in separate subpaths when they are large + +Good skill design keeps the rapid path clear without burying the model in unnecessary context. + +### 5. Add tests with realistic prompts + +Create 2–5 realistic eval prompts that match how a user would actually request this skill. + +For each prompt, define: +- the user request +- the expected behavior or output +- whether it is a deterministic or subjective evaluation + +Use a baseline comparison when comparing skill performance is worthwhile. + +### 6. Evaluate and iterate + +Run the skill against realistic prompts and compare to a baseline or previous version. + +Review: +- whether the skill triggered at the right times +- whether the outputs were usable and complete +- whether there were obvious failures or over-broad behavior +- whether the description should be more forceful or more precise + +Then revise the skill based on the results. + +### 7. Optimize for trigger quality + +If the skill is under-triggering or over-triggering, improve the `description` and `When to Use` wording. + +Best practice: +- include the task and trigger phrases explicitly +- mention the context where the skill is useful even if the user does not say the exact domain name +- keep the description crisp, but make it actionable enough that the agent will choose it when appropriate + +### 8. Finalize and document + +Before concluding, verify: +- the skill is narrow and purposeful +- boundaries are explicit +- outputs are concrete and consistent +- examples match the stated behavior +- the workflow is general enough to be reused + +## Output Format + +Use a clear skill brief like: + +```markdown +## Skill plan + +**Skill name**: +**Purpose**: +**Primary trigger phrases**: +**Output**: +**Key boundaries**: +**Eval prompts**: <2-5 prompts to validate behavior> +``` + +## Examples + +### Example Input + +```text +Turn this checklist workflow into a reusable skill for my team. It should help agents review project setup, identify missing dependencies, and produce a concise remediation plan. +``` + +### Example Output + +```markdown +## Skill plan + +**Skill name**: project-setup-reviewer +**Purpose**: Check whether a repo or project is configured correctly and produce a short remediation plan. +**Primary trigger phrases**: review project setup, find missing dependencies, audit local environment, identify setup gaps +**Output**: a checklist of missing pieces and a prioritized fix plan +**Key boundaries**: do not rewrite the codebase; do not invent credentials; do not assume deployment targets without evidence +**Eval prompts**: 1) audit this project setup, 2) find missing environment files, 3) identify the next steps to make this repo runnable +``` + +## Boundaries + +- Do not create a skill that is broader than a single, clear responsibility +- Do not leave ambiguous trigger conditions; the description should make selection obvious +- Do not write evals that are subjective unless the skill is inherently subjective +- Do not skip iteration; strong skills are refined against real prompts and feedback +- Do not hide dangerous or unauthorized actions inside a skill; keep scope safe and transparent + +## Authoritative references + +- Anthropic `anthropics/skills` — official `skill-creator` guidance and skill authoring patterns +- `agentskills/agentskills` — open standard for agent skill metadata and structure +- VRIL LABS `skill-jam` — strong examples of skill discovery and authoring discipline diff --git a/.github/skills/skill-navigator/SKILL.md b/.github/skills/skill-navigator/SKILL.md new file mode 100644 index 0000000..4bb1585 --- /dev/null +++ b/.github/skills/skill-navigator/SKILL.md @@ -0,0 +1,156 @@ +--- +name: skill-navigator +description: > + Find the right agent skill for a task, decide when to use a skill vs. inline reasoning, and sequence + skill calls efficiently. Use when the user asks which skill to use, when the workflow is uncertain, + or when a task would benefit from a proven reusable process. +--- + +# Skill Navigator + +Use this skill to find the best skill for the current task and choose the right invocation timing. The goal is to maximize quality and minimize wasted tokens. + +This is a distilled synthesis of the strongest skill-discovery patterns across the agent-skills ecosystem, especially the skill-discovery guidance in VRIL LABS `skill-jam` and the broader industry practice of matching tasks to explicit `When to Use` sections. + +## When to Use + +- Agent is unsure which skill applies to a task +- User asks "what skills are available?" or "which skill should I use for X?" +- A task has a known structure and a skill could encode it cleanly +- The task would benefit from a repeatable process or reference workflow +- A session is starting and the agent should orient to available capabilities +- A phase is complete and the next skill in the chain needs to be chosen + +## Process + +### 1. Decide whether a skill is needed + +Before searching, judge whether a skill is appropriate. + +Use a skill when: +- the task has a well-defined structure +- the workflow is repetitive or domain-specific +- external APIs or platform operations are involved +- the process would take multiple steps to re-derive +- output needs a consistent format + +Do not use a skill when: +- the task is a simple direct answer +- no known skill fits and the work is cheaper inline +- the skill would overcomplicate a small task +- the skill would duplicate work already done in the same session + +### 2. Start from the obvious category + +Check the likely skill families first: +- code review or debugging +- research and web search +- file and project analysis +- testing and validation +- deployment and operations +- documentation and writing +- domain-specific services or tools + +Prefer the most likely matches before scanning a broad list. + +### 3. Filter by description and trigger phrases + +Look for the best matches using these signals: +- task verb and noun +- technology or domain keywords +- expected output type +- the skill's explicit `When to Use` or description + +If multiple skills look plausible, narrow to the top 2–4 candidates by fit, not by volume. + +### 4. Read the candidate `When to Use` sections + +This is the authoritative decision gate. If the current task matches a bullet or trigger phrase, the skill is a valid fit. If not, do not recommend it. + +Use the rule: +- one clear match is enough +- more than 4 recommendations is usually too broad +- avoid recommending the same skill twice in a session for the same purpose + +### 5. Choose the right timing + +Invoke at session start when: +- the task requires repo reconnaissance before implementation +- the work depends on external research or tooling context +- the skill sets up a reliable process for the whole session + +Invoke mid-task when: +- the implementation or diagnosis has a clear stage boundary +- the user needs a focused specialist step +- a review or validation step is needed before finish + +Invoke at the end when: +- documentation or release readiness is needed +- final validation or cleanup is required + +### 6. Prefer chains when they add clarity + +When a task has multiple stages, use a small chain rather than a single "do everything" skill. + +A typical flow is: +- reconnaissance or research skill +- implementation or domain skill +- review, validation, or documentation + +Keep chains minimal and purposeful. + +### 7. Return a concise navigation recommendation + +Recommend the best fit with: +- task summary +- recommended skill(s) +- invocation timing +- brief rationale +- optional chain sequence + +## Output Format + +```markdown +## Skill Navigation + +**Task**: +**Recommended skill(s)**: +**Invocation timing**: +**Rationale**: <1–2 sentence explanation> +**Chain**: +``` + +## Examples + +### Example Input + +```text +We need to set up database migrations for a Node.js app. Which skill should we use first, and when? +``` + +### Example Output + +```markdown +## Skill Navigation + +**Task**: Add PostgreSQL migration workflow to a Node.js project +**Recommended skill(s)**: `codebase-recon`, `database-migration` +**Invocation timing**: Start with `codebase-recon` now, then use `database-migration` before implementing +**Rationale**: `codebase-recon` finds reference patterns and existing conventions first, while `database-migration` encodes the actual migration process and common failure modes. +**Chain**: `codebase-recon` → `database-migration` → implement → validate +``` + +## Boundaries + +- Do not recommend a skill unless its trigger conditions clearly match the task +- Do not recommend more than 4 skills for a single task +- Do not recommend a skill twice for the same work in one session +- Do not force a skill when inline reasoning is clearly cheaper and cleaner +- Always include invocation timing; "use skill X" by itself is incomplete +- If no skill fits, say so clearly and recommend inline reasoning + +## Authoritative references + +- VRIL LABS `skill-jam` `skill-navigator` and `skill-builder` guidance +- Anthropic `anthropics/skills` guidance on skill creation and metadata design +- `agentskills/agentskills` open standard for skills, descriptions, and discovery patterns