Skip to content

docs: clarify Kubernetes OOM and scheduling failures - #771

Merged
hongyi-chen merged 13 commits into
mainfrom
factory/k8s-resource-troubleshooting
Oct 7, 2026
Merged

hongyi-chen merged 13 commits into
mainfrom
factory/k8s-resource-troubleshooting

Conversation

@warp-agent-staging

@warp-agent-staging warp-agent-staging Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Clarifies how operators diagnose and size self-hosted Kubernetes tasks, so worker-daemon resources, task-container OOMs, and scheduler capacity failures aren't confused. Closes #769.

#770 has merged, so this PR now targets main directly. main also moved the self-hosting docs from platform/ to factories/. This update merges main, resolves the conflicts at the new paths, and re-verifies every claim against current source.

Changes

factories/self-hosting/troubleshooting.mdx

  • Replaces the dense Kubernetes failure list with symptom, verify, and fix paths for FailedScheduling, OOMKilled, eviction, and exit code 143.
  • Keeps SIGTERM (exit 143) separate from an OOM report, and failures before scheduling separate from failures in a running container.
  • Documents the unschedulableTimeout and activeDeadlineSeconds behavior operators run into.
  • Trims the metrics steps from seven to six without dropping the bind-address check.

factories/self-hosting/managed-kubernetes.mdx

  • Adds a "Size task containers" section covering worker.resources vs task resources, pod_template, and runner instance shapes (including precedence and the generated-init-container limit).
  • Fixes the unschedulableTimeout default and rewrites the capacity section.
  • Restores the base intro and preserves source-backed scaling, security-context, liveness, in-cluster-auth, and monitoring guidance.

factories/self-hosting/reference.mdx

  • Fixes the unschedulable_timeout default.

Corrections found while re-verifying

  • unschedulable_timeout / kubernetesBackend.unschedulableTimeout defaults to 10m, not 30s as main currently says (three places). Source: defaultUnschedulableFailureDelay in internal/worker/kubernetes.go, unschedulableTimeout: "10m" in the chart values.yaml, and the worker README.
  • The runners link now points to /factories/runners/.

Verified against source

warpdotdev/oz-agent-worker@e3ce13d (internal/worker/kubernetes.go, charts/oz-agent-worker/values.yaml, README.md):

  • Chart worker.resources requests 100m CPU and 128Mi memory, with no limits.
  • Task containers get no worker-defined CPU or memory defaults.
  • An instance shape sets request equal to limit per specified axis and overrides matching pod_template values, for the task container only.
  • The worker classifies OOMKilled, eviction, unschedulable, and deadline failures. Exit 143 is handled as SIGTERM guidance, not as an OOM.
  • activeDeadlineSeconds defaults to 28800 (eight hours). Failed Jobs are retained for 24 hours by default.

Intentional omissions

  • No universal task size or blanket overprovisioning advice, since the product defines no workload-independent baseline.
  • Internal retry counts and policies, which operators can't act on.

Content design plan

  • Reader and job: A Kubernetes operator diagnosing why a self-hosted task stopped or never started, then sizing the workload without changing unrelated workers.
  • Gap today: Troubleshooting combined resource failures, didn't cover OOMKilled or exit 143, and didn't separate daemon resources from task resources.
  • Change: Symptom-first OOM, eviction, SIGTERM, and scheduling paths plus concise sizing guidance. Excludes universal sizing numbers and internal retry behavior.

Verification

  • npm run build - passed.
  • check_links.py --internal-only - passed; 4,330 internal links checked, 0 broken.
  • style_lint.py --changed - passed; 3 changed files scanned, with 4 non-blocking warnings for bolded list lead-ins not in the glossary.
  • git diff --check - passed.
  • troubleshooting.mdx is 1,496 words, within the 1,500-word compression budget.
  • managed-kubernetes.mdx is 2,074 words, over the 1,500-word feature-doc budget. main was already at 1,982 words, and this PR adds focused resource guidance while cutting elsewhere. Splitting the page is out of scope.
  • Trunk CLI wasn't available, so trunk check was not run.

Unverified claims

None. The claims above were checked against the source listed.

Documentation risk

Risk: engineering-review-required
Rationale: Adds technical claims about self-hosted Kubernetes task resources, runner shapes, scheduling timeouts, and failure classification, and corrects a documented default.
Source files consulted: oz-agent-worker@e3ce13d: internal/worker/kubernetes.go, charts/oz-agent-worker/values.yaml, README.md
Docs override: none

Co-Authored-By: Oz oz-agent@warp.dev

warp-agent-staging Bot and others added 2 commits September 19, 2026 18:53
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
@cla-bot cla-bot Bot added the cla-signed label Sep 19, 2026
@vercel

vercel Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
docs Ready Ready Preview Oct 7, 2026 12:31am UTC

Request Review

@warp-agent-staging warp-agent-staging Bot added factory:docs-factory Label associated to the "docs-factory" factory warpy-factory Opened by the Warp factory agents labels Sep 19, 2026
@warp-agent-staging

Copy link
Copy Markdown
Contributor Author

This PR was generated with Warp.

Comment @warp-staging-factory on this PR to send it follow-up work.

View run View conversation View origin

Co-Authored-By: Oz <oz-agent@warp.dev>
warp-agent-staging Bot and others added 3 commits September 20, 2026 17:19
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
@warp-agent-staging
warp-agent-staging Bot changed the base branch from main to docs/self-hosted-kubernetes-troubleshooting September 20, 2026 17:28
Co-Authored-By: Oz <oz-agent@warp.dev>
warp-agent-staging Bot and others added 2 commits September 20, 2026 17:32
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Base automatically changed from docs/self-hosted-kubernetes-troubleshooting to main September 25, 2026 02:53
hongyi-chen and others added 2 commits October 6, 2026 22:25
…rnetes guidance

Resolve conflicts from main moving self-hosting docs from platform/ to
factories/. Reapply the OOM, scheduling, eviction, and exit-code-143
troubleshooting and the task sizing guidance at the new paths, and
re-verify them against oz-agent-worker.

- Fix unschedulable_timeout default (10m, not 30s) in managed-kubernetes
  and reference pages.
- Add a "Size task containers" section; link runners at /factories/runners/.
- Rewrite for plainer prose; remove duplicated and padded text.

Co-Authored-By: Warp <agent@warp.dev>
Co-Authored-By: Warp <agent@warp.dev>
@warp-for-oss

warp-for-oss Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

@hongyi-chen

I'm starting a first review of this pull request.

You can view the conversation on Warp.

I completed the review and no human review was requested for this pull request.

Comment /warp-agent-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

The independent agent completed its review for this commit.

Findings

  • src/content/docs/factories/self-hosting/troubleshooting.mdx:102 — [IMPORTANT] This page was within the 1500-word feature-doc budget on main and is now 1523 words (check_compression_contract.py exits 1). The PR body justifies the overage only for managed-kubernetes.mdx. Requested change: do a deletion-only pass to bring the page back to 1500 words or fewer (for example, trim the new Verify and Fix prose in the Kubernetes task-failure subsections), or add a reason for this page's overage to the PR body.
  • src/content/docs/factories/self-hosting/managed-kubernetes.mdx:286 — [SUGGESTION] The removed Operational notes section was described as restated elsewhere, but its security-context fact (non-root runAsUser: 10001, allowPrivilegeEscalation: false, all capabilities dropped) now appears nowhere in src/content/docs. The chart values.yaml still sets these. Requested change: restore one sentence about the Deployment's default security context, for example under 'What the chart deploys', or confirm the removal is intentional.
  • src/content/docs/factories/self-hosting/troubleshooting.mdx:112 — [SUGGESTION] The 'Fix (all backends)' step 'Ensure the worker machine or cluster has sufficient resources (CPU, memory, disk)' was deleted, so the Docker and Direct task-failure sections no longer mention host resource exhaustion. The new Kubernetes subsections cover this only for Kubernetes. Requested change: restore a one-line resource check under ### Docker backend (task failures) and the Direct backend section.
  • src/content/docs/factories/self-hosting/managed-kubernetes.mdx:199 — [SUGGESTION] The PR is framed as a cut, but the page word count rose from 1982 on main to 2035, still well over the 1500-word budget. The overage is justified in the PR body, but the new ## Size task containers section repeats guidance that the troubleshooting page and the worker.resources bullet (line 114) already give, including the OOMKilled and FailedScheduling pointer at line 210. Requested change: remove the sentence at line 210 that points to the troubleshooting page, or shorten the section so the page ends up smaller than on main.

Verdict

Request changes

@warp-for-oss warp-for-oss Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overview

This PR updates the self-hosted Kubernetes documentation to separate worker Deployment resources from task-container resources, correct the unschedulable timeout default, and expand troubleshooting for scheduling, OOM, eviction, and SIGTERM cases.

Concerns

  • No blocking correctness, security, style, link, or comment/test-quality concerns were found in the attached diff.
  • spec_context.md states that no approved or repository spec context was found, so there is no material spec drift to report.
  • The PR body's engineering-review-required documentation risk classification is appropriate for the changed technical claims and includes the source files consulted.

Verdict

Found: 0 critical, 0 important, 0 suggestions

Approve

Comment /warp-agent-review on this pull request to retrigger a review (up to 3 times on the same pull request).

Powered by Oz

Comment thread src/content/docs/factories/self-hosting/managed-kubernetes.mdx Outdated
Comment thread src/content/docs/factories/self-hosting/managed-kubernetes.mdx
Comment thread src/content/docs/factories/self-hosting/troubleshooting.mdx
@hongyi-chen
hongyi-chen merged commit ce3dd39 into main Oct 7, 2026
15 of 16 checks passed
@hongyi-chen
hongyi-chen deleted the factory/k8s-resource-troubleshooting branch October 7, 2026 06:12

This branch was successfully deployed

1 active deployment
Preview — 661a23da Deployed Oct 7, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cla-signed factory:docs-factory Label associated to the "docs-factory" factory warpy-factory Opened by the Warp factory agents

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Document self-hosted Kubernetes OOM vs scheduling failures

2 participants