Repository navigation
docs: clarify Kubernetes OOM and scheduling failures - #771
Conversation
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
This PR was generated with Warp. Comment |
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
Co-Authored-By: Oz <oz-agent@warp.dev>
…rnetes guidance Resolve conflicts from main moving self-hosting docs from platform/ to factories/. Reapply the OOM, scheduling, eviction, and exit-code-143 troubleshooting and the task sizing guidance at the new paths, and re-verify them against oz-agent-worker. - Fix unschedulable_timeout default (10m, not 30s) in managed-kubernetes and reference pages. - Add a "Size task containers" section; link runners at /factories/runners/. - Rewrite for plainer prose; remove duplicated and padded text. Co-Authored-By: Warp <agent@warp.dev>
Co-Authored-By: Warp <agent@warp.dev>
|
I'm starting a first review of this pull request. You can view the conversation on Warp. I completed the review and no human review was requested for this pull request. Comment Powered by Oz |
There was a problem hiding this comment.
Review summary
The independent agent completed its review for this commit.
Findings
src/content/docs/factories/self-hosting/troubleshooting.mdx:102— [IMPORTANT] This page was within the 1500-word feature-doc budget on main and is now 1523 words (check_compression_contract.py exits 1). The PR body justifies the overage only formanaged-kubernetes.mdx. Requested change: do a deletion-only pass to bring the page back to 1500 words or fewer (for example, trim the new Verify and Fix prose in the Kubernetes task-failure subsections), or add a reason for this page's overage to the PR body.src/content/docs/factories/self-hosting/managed-kubernetes.mdx:286— [SUGGESTION] The removedOperational notessection was described as restated elsewhere, but its security-context fact (non-rootrunAsUser: 10001,allowPrivilegeEscalation: false, all capabilities dropped) now appears nowhere insrc/content/docs. The chartvalues.yamlstill sets these. Requested change: restore one sentence about the Deployment's default security context, for example under 'What the chart deploys', or confirm the removal is intentional.src/content/docs/factories/self-hosting/troubleshooting.mdx:112— [SUGGESTION] The 'Fix (all backends)' step 'Ensure the worker machine or cluster has sufficient resources (CPU, memory, disk)' was deleted, so the Docker and Direct task-failure sections no longer mention host resource exhaustion. The new Kubernetes subsections cover this only for Kubernetes. Requested change: restore a one-line resource check under### Docker backend (task failures)and the Direct backend section.src/content/docs/factories/self-hosting/managed-kubernetes.mdx:199— [SUGGESTION] The PR is framed as a cut, but the page word count rose from 1982 on main to 2035, still well over the 1500-word budget. The overage is justified in the PR body, but the new## Size task containerssection repeats guidance that the troubleshooting page and theworker.resourcesbullet (line 114) already give, including the OOMKilled and FailedScheduling pointer at line 210. Requested change: remove the sentence at line 210 that points to the troubleshooting page, or shorten the section so the page ends up smaller than on main.
Verdict
Request changes
There was a problem hiding this comment.
Overview
This PR updates the self-hosted Kubernetes documentation to separate worker Deployment resources from task-container resources, correct the unschedulable timeout default, and expand troubleshooting for scheduling, OOM, eviction, and SIGTERM cases.
Concerns
- No blocking correctness, security, style, link, or comment/test-quality concerns were found in the attached diff.
spec_context.mdstates that no approved or repository spec context was found, so there is no material spec drift to report.- The PR body's
engineering-review-requireddocumentation risk classification is appropriate for the changed technical claims and includes the source files consulted.
Verdict
Found: 0 critical, 0 important, 0 suggestions
Approve
Comment /warp-agent-review on this pull request to retrigger a review (up to 3 times on the same pull request).
Powered by Oz



Summary
Clarifies how operators diagnose and size self-hosted Kubernetes tasks, so worker-daemon resources, task-container OOMs, and scheduler capacity failures aren't confused. Closes #769.
#770 has merged, so this PR now targets
maindirectly.mainalso moved the self-hosting docs fromplatform/tofactories/. This update mergesmain, resolves the conflicts at the new paths, and re-verifies every claim against current source.Changes
factories/self-hosting/troubleshooting.mdxFailedScheduling,OOMKilled, eviction, and exit code143.SIGTERM(exit143) separate from an OOM report, and failures before scheduling separate from failures in a running container.unschedulableTimeoutandactiveDeadlineSecondsbehavior operators run into.factories/self-hosting/managed-kubernetes.mdxworker.resourcesvs task resources,pod_template, and runner instance shapes (including precedence and the generated-init-container limit).unschedulableTimeoutdefault and rewrites the capacity section.factories/self-hosting/reference.mdxunschedulable_timeoutdefault.Corrections found while re-verifying
unschedulable_timeout/kubernetesBackend.unschedulableTimeoutdefaults to10m, not30sasmaincurrently says (three places). Source:defaultUnschedulableFailureDelayininternal/worker/kubernetes.go,unschedulableTimeout: "10m"in the chartvalues.yaml, and the worker README./factories/runners/.Verified against source
warpdotdev/oz-agent-worker@e3ce13d(internal/worker/kubernetes.go,charts/oz-agent-worker/values.yaml,README.md):worker.resourcesrequests100mCPU and128Mimemory, with no limits.pod_templatevalues, for thetaskcontainer only.OOMKilled, eviction, unschedulable, and deadline failures. Exit143is handled asSIGTERMguidance, not as an OOM.activeDeadlineSecondsdefaults to 28800 (eight hours). Failed Jobs are retained for 24 hours by default.Intentional omissions
Content design plan
OOMKilledor exit143, and didn't separate daemon resources from task resources.Verification
npm run build- passed.check_links.py --internal-only- passed; 4,330 internal links checked, 0 broken.style_lint.py --changed- passed; 3 changed files scanned, with 4 non-blocking warnings for bolded list lead-ins not in the glossary.git diff --check- passed.troubleshooting.mdxis 1,496 words, within the 1,500-word compression budget.managed-kubernetes.mdxis 2,074 words, over the 1,500-word feature-doc budget.mainwas already at 1,982 words, and this PR adds focused resource guidance while cutting elsewhere. Splitting the page is out of scope.trunk checkwas not run.Unverified claims
None. The claims above were checked against the source listed.
Documentation risk
Risk: engineering-review-required
Rationale: Adds technical claims about self-hosted Kubernetes task resources, runner shapes, scheduling timeouts, and failure classification, and corrects a documented default.
Source files consulted: oz-agent-worker@e3ce13d: internal/worker/kubernetes.go, charts/oz-agent-worker/values.yaml, README.md
Docs override: none
Co-Authored-By: Oz oz-agent@warp.dev