Skip to content

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ppstats

Statistical functions and distributions, in POST Python.

ppstats reimplements scipy.stats in POST Python — every kernel is fully-typed Python that runs under the standard CPython interpreter and compiles ahead-of-time to native code (a plain C shared library and a NumPy ufunc extension module) with the POST Python reference compiler.

Status: Active — descriptive reductions and six continuous distribution families are verified in interpreted, native-library, and NumPy-ufunc modes. Part of the PostSciPy effort to rebuild SciPy one subpackage at a time as the compiler's proving ground.

Implemented functions

Family Module Functions
Descriptive _descriptive mean, variance, gmean, hmean, moment(a, order), skew, kurtosis, sem(a, ddof), zscore(a, ddof)
Continuous distributions _distributions <name>_pdf, <name>_cdf, <name>_ppf for norm, logistic, expon, uniform, laplace, and cauchy

Descriptive functions are reduction gufuncs ((n)->(), (n),()->(), or (n),()->(n)) that broadcast over batch dimensions. Numeric scipy parameters are positional inputs; boolean scipy modes are fixed to their scipy defaults (skew biased, kurtosis Fisher + biased). mean/variance mirror numpy.mean/numpy.var and are exposed as documented conveniences. Distribution functions are scalar ufuncs.

import numpy as np
from ppstats import skew, sem, zscore

skew([1.0, 2.0, 3.0, 4.0, 10.0])          # 1.1384199576606167
sem([1.0, 2.0, 3.0, 4.0, 5.0], 1)         # 0.7071067811865476  (ddof=1, scipy default)
zscore(np.array([[1., 2., 3.], [4., 6., 8.]]), 0)   # batches broadcast

Distribution loc and scale parameters are explicit positional inputs and broadcast like the value input:

from ppstats import norm_cdf, logistic_ppf

norm_cdf(1.5, 0.5, 2.0)       # 0.69146247... (x, loc, scale)
logistic_ppf(0.8, 0.5, 2.0)   # 3.27258872... (q, loc, scale)

Normal CDF/PPF accuracy is inherited from ppspecial, whose ndtr/ndtri error bound is absolute, not relative: measured against SciPy 1.18 on a dense grid, <1.2e-7 for the CDF and ~2.5e-7 * scale for the PPF. Reference cases pass at 2e-7 relative, but near norm_ppf's zero crossing (q = ndtr(-loc/scale)) relative error is unbounded — compare with atol + rtol*|ref|, never pure rtol. Other distribution kernels match SciPy references to 1e-12 relative on the test grid. scale > 0 is the valid domain. Invalid-scale behavior is deliberately pinned but is not yet SciPy-compatible; finite out-of-range PPF q values temporarily fold to the corresponding endpoint result. Valid unbounded PPF endpoints use exact ±1e308 sentinels until POST Python can lower IEEE infinity constants. NaN inputs propagate, including through PPF endpoint branches; PDFs tend to 0 and CDFs to 0/1 at infinite x. See docs/issues/ppstats-distribution-invalid-domain.md for the full pinned matrix and #36 follow-up.

NumPy is an optional extra (ppstats[numpy]), but interpreted execution of the descriptive reduction gufuncs currently requires it: postpyc 0.3.0's no-numpy fallback calls gufunc kernels without their output buffer (reproducer archived at docs/issues/postpyc-nonumpy-gufunc-fallback.md). The scalar distribution ufuncs run interpreted without numpy.

Building

pixi install -e dev
pixi run -e dev test            # test suite (interpreted + compiled modes)
pixi run -e dev build-dist      # pure Python wheel and source distribution
pixi run -e dev build-native    # plain C shared library + header + manifest
pixi run -e dev build-prefix    # libppstats layout under dist/prefix
pixi run -e dev build-ext       # ppstats_native, a NumPy-ufunc extension

Hatchling builds the Python distribution as a pure wheel and source archive; the expected wheel tag is py3-none-any. The native library and extension are explicit, separate build outputs, and installing ppstats never compiles native code.

Baseline accuracy report

pixi run -e dev accuracy-report       # hardcoded references, interpreted mode
pixi run -e dev accuracy-report-all   # build-ext, then hardcoded interpreted + native

# From an environment that already provides NumPy and SciPy:
python scripts/accuracy_report.py --reference scipy --mode interpreted

The report covers one canonical baseline case per public kernel; it is not the final domain-wide accuracy table. Native report checks establish only that the extension is local/source-current by path and mtime. Use a fresh build-ext (the accuracy-report-all dependency) for authoritative validation. This reporting layer does not change Hatch or native packaging behavior.

When ppstats_native is importable, import ppstats prefers the compiled gufuncs automatically (ppstats.__native_available__).

Primary compiler pressure this package generates: cross-package POST dependencies (ppspecial), reduction gufuncs.

Start here

  1. Read the POST Python spec and the PostSciPy roadmap (package map, working rules, capability matrix).
  2. Copy the layout of ppspecial, the exemplar package: ppstats/ sources, tests/, scripts/build_native.py, scripts/build_ext.py, a pixi workspace with test / build-native / build-ext tasks, a git dependency on postpython, and a ROADMAP.md tracking targets and upstream requests.
  3. Start with a slice from "Compiles today" below; land it as a small PR with tests in both execution modes.

First slices

Compiles today

  • Descriptive reductions as gufuncs (n)->(): mean, variance, skew, kurtosis, moment, gmean, hmean, sem; zscore as (n)->(n) Done (ROADMAP Target 1)
  • Distribution kernels built on ppspecial (the first cross-package POST dependency): norm, logistic, expon, uniform, laplace, and cauchy pdf/cdf/ppf Done (ROADMAP Target 3)

Blocked on compiler capabilities

File these as postpython issues with minimal reproducers when you start on them — the filing is part of the work and drives the compiler roadmap.

  • rvs sampling — needs a POST RNG model (design discussion in postpython first)
  • fit — needs ppoptimize / callable parameters
  • rv_continuous-style class framework — needs structs

Working rules (summary)

  • Pure POST Python: no compiler-specific escape hatches; every kernel runs interpreted and compiled.
  • scipy is the reference, never a runtime dependency. Tests may use it optionally; prefer deterministic hardcoded reference values.
  • Compiler gaps go upstream as postpython issues with reproducers, not silent workarounds.
  • Verify against released postpyc by default; use postpython main when chasing an unreleased compiler feature.
  • Document accuracy targets and reference sources per function.

The full rules and the definition of done live in the PostSciPy roadmap.

About

Statistical functions and distributions, in POST Python. POST Python rebuild of scipy.stats (PostSciPy effort).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages