Skip to content

Latest commit

 

History

History
109 lines (66 loc) · 5.86 KB

File metadata and controls

109 lines (66 loc) · 5.86 KB

Live switch architecture and performance

Live switch is an explicit CPython 3.13 optimization for repeated heterogeneous routing inside one active function call. It is not the default switch implementation and is not universally faster than portable lowering.

Execution model

Portable switch converts supported source shapes into normal non-self-modifying Python bytecode and table data.

Live switch instead keeps route bodies in-frame and maintains a verified jump gate. Each dispatch resolves a route and updates that gate so execution jumps directly to the selected body.

Because this depends on CPython code-object/adaptive-buffer layout, live mode is subject to stricter runtime and concurrency checks than portable mode.

Public modes

enable_switch(mode=...) supports:

  • auto — portable by default; an explicit live_threshold= may make eligible general plans live;
  • portable — force non-self-modifying lowering;
  • isolated — cached live clone per thread/active depth;
  • per_call — fresh live clone for each invocation;
  • thread_local — one live clone per thread, not re-entry safe;
  • fast — shared live code object for a proven single-active-call environment.

Do not choose a mode from its name alone. Match it to the workload and re-entry/concurrency contract.

Plan-aware auto selection

An explicit live_threshold= is not a route-count-only switch. The compiler first evaluates the portable plan. Direct-value, expression-template, and statement-template plans can remain portable even above the threshold because they already collapse to compact table/template code.

Large general/balanced plans can remain candidates for live mode, subject to runtime checks.

This prevents a strong portable representation from being replaced simply because a function has many cases.

Live engines

Explicit live modes accept live_engine=:

  • auto — prefer the optional native dispatcher after self-test, otherwise use the ctypes implementation;
  • native — require _livegate and fail if it is unavailable/unsupported;
  • ctypes — use the Python/ctypes gate path, mainly for fallback and comparison.

The native accelerator fuses route lookup and gate update. It is an optimization layer, not a semantic requirement.

Runtime checks

Before live mutation is accepted, the package validates the relevant CPython layout and reversible gate-write assumptions. runtime_diagnostics(full=True) exposes live layout/native status.

Core portable switch qualification occurs during root import; mutation-sensitive live checks remain lazy.

Free-threaded CPython

Live switch is unsupported on free-threaded CPython 3.13. The native extension is not imported as a way to silently re-enable GIL behavior. Use portable switch.

Workload fit

Strong live candidates typically have all of these characteristics:

  • one outer call performs many dispatches;
  • route bodies are heterogeneous enough that portable template lowering cannot collapse them well;
  • dispatch cost is a meaningful part of total work;
  • the live mode's re-entry/concurrency contract is acceptable.

Dense VM/interpreter loops and some parser/state-machine dispatchers fit this profile.

Prefer portable mode for common cases such as:

  • direct-value routing that already approaches dictionary lookup;
  • expression/statement-template-compatible routes;
  • HTTP/RPC routing with one router call per request;
  • short async handlers;
  • heavy route bodies where dispatch is a small fraction of total work;
  • shared/reentrant code whose live isolation cost would dominate.

Performance evidence

The canonical live-mode evidence is retained under benchmarks/results/BENCHMARK_LIVE_EXTENSIVE_V122.*; v122 is an engineering evidence identifier, not a package version.

Its key conclusion is workload-dependent: recorded dense VM/parser cases benefit substantially from native live mode, while HTTP/direct/template-friendly and some sparse/heavy-server controls tie or favor portable mode. A separate per-request control demonstrates that a busy server does not automatically make live routing profitable when each request enters the router only once.

See Benchmarks for exact tables, methodology, and reproduction commands. Exact numbers live there rather than being duplicated in this architecture guide.

Server, async, and recursion guidance

The outer-call structure matters more than a generic notion of a “hot server.” If each request/coroutine invokes the router once or only a few times, clone/isolation and gate overhead may dominate.

Use portable mode by default for async/shared server code. If a real workload performs many internal dispatches per invocation and live mode measures better, use the safest live mode that satisfies the application's re-entry/concurrency requirements.

Code size and compile cost

Portable template lowering can reduce very large case-oriented source to compact executable bytecode plus table data. Live mode retains route bodies inline, so it can cost more to compile and occupy more code space.

For a dense VM that executes many internal dispatches, that cost may amortize quickly. For compact direct/template routing, it may never amortize. Benchmark decoration/import cost as well as steady-state runtime.

Reproducing evidence

Use the benchmark hierarchy in benchmarks/README.md:

# Establish whether extension support helps the workload.
python benchmarks/scripts/benchmark_primary_v130.py --quick

# Then compare portable/live implementations on live-appropriate workloads.
python benchmarks/scripts/benchmark_switch_live_dispatch_v121.py
python benchmarks/scripts/benchmark_live_workloads_v122.py
python benchmarks/scripts/benchmark_live_server_modes_v122.py
python benchmarks/scripts/benchmark_live_server_per_request_v122.py

Performance changes to live dispatch should be evaluated against the cross-workload matrix, not only a dispatch-core microbenchmark.