Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 8 additions & 63 deletions docs/backends/nanvix/nanvix.md
Original file line number Diff line number Diff line change
Expand Up @@ -181,70 +181,15 @@ including egress allow with omitted ingress defaults.
Disabled networking prevents guest socket creation (`OSError: [Errno 134]`).
Unrestricted networking includes host-backed bind/listen capabilities;
NanVix cannot independently enforce ingress or host-loopback restrictions.
Directional egress rules are explicitly rejected because the legacy IPv4
filter does not implement their full semantics, including default-deny DNS.
Directional egress rules are explicitly rejected rather than translated to
the guest's IPv4 host filter, which cannot implement their full semantics,
including default-deny DNS.
Runtime proxy configuration is also unsupported.

### Legacy per-host filter implementation (compatibility/reference only)

The following describes the retained legacy runtime filter, not accepted v1.1
JSON vocabulary. The exact cutover does not silently translate directional
rules into this weaker contract.

Legacy `defaultPolicy` and host-list interactions follow the
[backend-agnostic network policy semantics](../../schema.md#legacy-network-host-list-semantics).
Invalid legacy combinations are rejected by shared policy validation:
`blockedHosts` requires an `allowedHosts` exception set under a block default,
and `allowedHosts` cannot be used under an allow default.

NanVix forwards the validated host list to the guest's host-side socket proxy,
which enforces egress at `connect()`. The guest filter is **allow-XOR-block**,
so NanVix rejects requests that supply both lists instead of dropping either
one.

Entries may be IPv4 literals (`93.184.216.34`), IPv4 CIDR blocks
(`10.0.0.0/8`), or hostnames. Hostnames are resolved to their IPv4 (A-record)
addresses at preflight; IPv6 (AAAA) results are dropped because the guest filter
is IPv4-only. Resolution failures are handled per direction so neither list ever
fails open:

- **allowlist** (deny-by-default): each dropped entry is logged as a warning and
the run continues, since dropping an entry only *narrows* access. If the list
resolves to **no** IPv4 address at all, the run is rejected at preflight rather
than silently allowing all traffic.
- **blocklist** (allow-by-default): **any** entry that resolves to no IPv4
address rejects the run at preflight. Silently dropping a blocked host would
let traffic the policy explicitly blocks flow freely, and the static preflight
filter cannot enforce a name that does not resolve β€” so the blocklist
fails closed.

**DNS:** in allowlist mode the guest daemon automatically exempts the DNS port
(53), so name resolution works without adding the resolver to `allowedHosts`.

Network proxies (`network.proxy`) are not supported and are rejected at
preflight.

```jsonc
{
"containment": "microvm",
"process": { "commandLine": "import urllib.request; ..." },
// Historical legacy shape, not accepted by the exact v1.1 contract:
"network": { "defaultPolicy": "allow" }
}
```

```jsonc
{
"containment": "microvm",
"process": { "commandLine": "import urllib.request; ..." },
// Historical legacy shape, not accepted by the exact v1.1 contract:
"network": { "allowedHosts": ["example.com", "10.0.0.0/8"] }
}
```

## Not Supported

| Workload | Error |
| ------------------------------- | ----------------------------------- |
| Both `allowedHosts` + `blockedHosts` | Rejected at preflight (mutually exclusive) |
| File writing outside `/mnt/rw/` | `OSError: Read-only file system` |
| Workload | Error |
| -------------------------------------------- | -------------------------------- |
| Mixed directional networking or egress rules | Rejected before VM creation |
| Runtime proxy | Rejected before VM creation |
| File writing outside `/mnt/rw/` | `OSError: Read-only file system` |
7 changes: 5 additions & 2 deletions docs/backends/process-container/networking.md
Original file line number Diff line number Diff line change
Expand Up @@ -205,7 +205,7 @@ middle rows provide different protections and are not ordered relative to each o

For identity-scoped proxies, the scoped peer rule and `privateNetworkClientServer` do not bypass Windows
Firewall's block-inbound-to-non-allowed-apps policy. A packaged AppContainer proxy uses the package-owned firewall
declaration shown in the [historical schema 0.8 manifest example](examples/0.8.0-schema.md); its application entry uses
declaration shown in the [proxy package manifest example](examples/0.8.0-schema.md); its application entry uses
`uap10:RuntimeBehavior="packagedClassicApp"` with `uap10:TrustLevel="appContainer"`. An unpackaged AppContainer proxy
requires its installer or administrator to own an equivalent rule scoped to the AppContainer profile SID, proxy
executable, and configured port.
Expand Down Expand Up @@ -279,7 +279,10 @@ proxy peer identity, and host-loopback allow fail with a typed unsupported-polic
## 3. WFP enforcement

PSEC applies outbound WFP filters in the OS's elevated context and owns their lifetime. AppContainer fallback does not
install directional WFP filters.
install directional WFP filters. It uses capabilities for supported direction
defaults and keeps the external runtime proxy setup. Its network audit records retain the firewall fields with
`firewall_rules_created` and `firewall_rules_removed` at `0`, `firewall_applied` at `false`, and
`firewall_removal_ok` at `true`.

WFP implements `egress` rules for public and private destinations. `internetClient` enables public-network access.
`privateNetworkClientServer`, selected through `ingress.default`, is the prerequisite for private-network access and
Expand Down
6 changes: 3 additions & 3 deletions docs/backends/windows-sandbox/windows-sandbox.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,9 +151,9 @@ non-default values.
| `filesystem.readwritePaths` | Existing directories mapped read-write at the same path |
| `filesystem.readonlyPaths` | Existing directories mapped read-only at the same path |
| `filesystem.deniedPaths` | Accepted outside shares; rejected when overlapping a mapped share |
| Default network policy `block` | Enforced by the guest firewall |
| Default network policy `allow` | Rejected |
| `allowedHosts` / `blockedHosts` | Rejected |
| Omitted network policy | Guest firewall blocks external networking |
| `network.egress` or `network.ingress` supplied (even empty or deny-only) | Rejected; omit both sections to use guest isolation |
| Retired `defaultPolicy` and host lists | Rejected by the exact contract |
| Network proxy | Rejected |

Mapped paths must be absolute existing directories. Files, nested mapped roots,
Expand Down
5 changes: 2 additions & 3 deletions docs/backends/wslc/wsl-container-getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -422,11 +422,10 @@ rules, but a WSLC container runs **without** `CAP_NET_ADMIN` (the SDK's
`Privileged` flag does not grant it), so those rules cannot be applied β€” and MXC
has no VM-level enforcement hook either (WSLC cannot expose one without breaking
other security promises such as MDE). Rather than fail the run at exec time,
such configs are **rejected at config-parse time**:
such configs are **rejected before provisioning**:

```
WSLc: per-host egress filtering (allowedHosts with defaultPolicy='block', or
blockedHosts with defaultPolicy='allow') is not supported. ...
WSLc does not support network.egress allow/deny rules; networking is all-or-nothing
```

Use a runtime proxy with unrestricted bridged networking for cooperative host
Expand Down
7 changes: 6 additions & 1 deletion docs/development/architecture/telemetry.md
Original file line number Diff line number Diff line change
Expand Up @@ -547,6 +547,11 @@ rejection records from the same invocation. A successful launch emits no
| `mxc.SandboxTornDown` | Per-run resources released, once per handle | ProcessContainer: `backend`, `identity`, `tier`, `pid`, `status`, `firewall_rules_removed`, `firewall_removal_ok`, `bfs_removed`, `proxy_stopped`, `preserve_policy`, `container_released`, `skip_reason`. IsolationSession: `backend`, `identity`, `phase`, `status`, `session_stopped`, `agent_user_deprovisioned`, `client_unregistered` |
| `mxc.ConfigRejected` | A request was refused before it could run | `correlation_id`, `backend`, `reason`, `offending_field`, `phase` |

For the ProcessContainer AppContainer fallback, the firewall fields remain in
these records for compatibility: no local firewall rules are created or
removed, `firewall_applied` is `false`, and `firewall_removal_ok` is `true`.
Proxy setup failures still set `mxc.NetworkPolicyApplied.status` to failure.

### Error semantics: `FallbackError` vs `ActivityError`

Tier selection can *degrade* (proceed with weaker enforcement) or *fail*
Expand Down Expand Up @@ -710,7 +715,7 @@ that it was skipped.
| Process outcome (M-ETW-1) | βœ… | βœ… | βœ… | βœ… (shared `create_process`) |
| Enforcement degradation (M-ETW-2) | βœ… (shared dispatcher; records the tier actually selected) | βœ… | n/a β€” no tier/fallback ladder exists for this backend | n/a |
| Policy hash (M-ETW-3) | βœ… | βœ… | βœ… | βœ… |
| Network policy (M-ETW-4) | βœ… (`enforcement_mode: capabilities` β€” policy travels in the sandbox spec and the OS enforces it, so `firewall_rules_created` is honestly `0`) | βœ… (`firewall` / `both`) | n/a β€” MXC rejects network and proxy policy for this backend before provisioning | n/a |
| Network policy (M-ETW-4) | βœ… (`enforcement_mode: capabilities` β€” policy travels in the sandbox spec and the OS enforces it, so `firewall_rules_created` is honestly `0`) | βœ… (`enforcement_mode: capabilities` for supported directional requests; egress default is reported as `allow` or `block`) | n/a β€” MXC rejects network and proxy policy for this backend before provisioning | n/a |
| Sandbox teardown (M-ETW-5) | βœ… | βœ… | βœ… | βœ… (`stop` and `deprovision` phases) |
| IsolationSession telemetry (M-ETW-6) | n/a | n/a | βœ… Applicable lifecycle events use `Microsoft.MXC`; no separate OS provider is assumed | βœ… Same provider path |
| Configuration rejection (M-ETW-7) | βœ… | βœ… | βœ… | βœ… (`phase` names the rejecting phase) |
Expand Down
2 changes: 1 addition & 1 deletion docs/development/plans/linux-wsl-roadmap-june-2026.md
Original file line number Diff line number Diff line change
Expand Up @@ -503,7 +503,7 @@ These items depend on the WSLC SDK team and are not unilaterally schedulable.

> **Why network enforcement must be container-scoped (host vs. VM vs. container).** Network policy can be enforced at three layers: the Windows **host** (Windows Firewall), the WSL2 **VM**, or the **container** network namespace inside the VM. GA decision **D6 (per-sandbox scoping)** requires every sandbox's policy to be independent β€” concurrent WSLC containers must not affect each other's access β€” and names the container network namespace as WSLC's scoping identity. A machine-wide **host** firewall can't attribute traffic to one container vs. another, so it violates D6 (and per **D8**, host firewalls apply *on top of* enforcement, never *as* it). A **VM-wide** rule fails the same way when one utility VM hosts multiple containers β€” sandbox A's rules would bleed into sandbox B. Only the **container namespace** is inherently per-sandbox, which is why it's the required enforcement point. The catch: MXC can't install rules into that namespace today (`Privileged` doesn't grant `CAP_NET_ADMIN`, and the VM may lack iptables tooling). Hence SDK dep #1 β€” a VM-level API that applies rules **scoped to a specific container's namespace**: physically enforced at the VM boundary, logically attributed to one container.
>
> **Contrast with Hyperlight/Nanvix, and the state-aware wrinkle.** Hyperlight (host-proxied sockets, per-instance) and Nanvix (per-guest egress filter) get D6 scoping for free because each sandbox *is* its own VM instance/process β€” no shared surface to bleed across. WSLC today is also effectively 1 sandbox : 1 VM (the one-shot flow creates a session, one container, then tears it down), but the highest-value WSLC optimization β€” **state-aware session reuse** (Misc #29), keeping a warm VM to amortize startup cost β€” makes one VM host **multiple** containers, at which point a host- or VM-wide rule genuinely bleeds across co-resident sandboxes. That is exactly when namespace-scoped enforcement (SDK dep #1) stops being merely cleaner and becomes mandatory.
> **Contrast with Hyperlight/Nanvix, and the state-aware wrinkle.** Hyperlight (network disabled for supported requests, per-instance) and Nanvix (all-deny or unrestricted networking, per-guest) keep their network posture scoped to each VM instance/process β€” no shared surface to bleed across. WSLC today is also effectively 1 sandbox : 1 VM (the one-shot flow creates a session, one container, then tears it down), but the highest-value WSLC optimization β€” **state-aware session reuse** (Misc #29), keeping a warm VM to amortize startup cost β€” makes one VM host **multiple** containers, at which point a host- or VM-wide rule genuinely bleeds across co-resident sandboxes. That is exactly when namespace-scoped enforcement (SDK dep #1) stops being merely cleaner and becomes mandatory.

---

Expand Down
Loading