Integer data types lack special values for -inf, inf and NaN. Especially
NaN as an indication for missing data would be useful in many scientific contexts.
Of course there is numpy.ma.MaskedArray around for the very same reason.
Nevertheless, it might sometimes be annoying to carry a separate mask array
around — and masked arrays are notoriously slow. In those cases, using a set of
numpy-compatible functions for the same job will do just fine.
This package provides such an implementation for a set of standard numpy
functions, treating integer arrays in such a way, that a designated sentinel
value resembles NaN:
- For signed integer types, the lowest negative value (
np.iinfo(dtype).min) is used as the missing value, e.g.-2147483648forint32. Large negative values are chosen deliberately, so that python indexing (from the end) is unlikely to run into them accidentally. - For unsigned integer types, the value
0is used. - For float and string types, the corresponding
NaN, empty bytes or empty string values are used. - For object arrays,
Noneis used.
pip install intnanor with uv:
uv pip install intnanThe package requires Python 3.12+ and works with numpy 2.4+. All public
functions are fully type annotated using numpy.typing.
Simply import the package and use the provided functions like their numpy
counterparts:
import numpy as np
import intnan
a = np.array([1, -(2**31), 3], dtype=np.int32) # -(2**31) marks a missing value
intnan.isnan(a) # array([False, True, False])
intnan.nansum(a) # 4
intnan.nanmean(a) # 2.0
intnan.fix_invalid(a) # array([1, 0, 3], dtype=int32)The following functions are provided by intnan. Where applicable, their
semantics mirror the corresponding numpy function, with missing values
ignored instead of propagated. All reductions accept the usual axis and
keepdims arguments like their numpy counterparts. All-missing slices
yield the missing value for integer dtypes and NaN for floating dtypes;
the arg* functions raise on all-missing slices like numpy does.
Missing value handling:
nanval(x)— return the missing value for a given array or data typeisnan(x)— boolean mask of missing values, works on arrays and scalarsfix_invalid(x, copy=True, fill_value=0)— replace missing valuesasfloat(x)— convert to a float array, missing values becomeNaNasint(x)— convert to an integer array, missing values become the integer missing valueanynan(x, axis=None),allnan(x, axis=None)— test for the presence of missing valuesnancount(x, axis=None)— number of valid values
Reductions:
nanmax(x),nanmin(x),nanptp(x)and their index counterpartsnanargmax(x),nanargmin(x)nansum(x),nanprod(x),nancumsum(x),nancumprod(x)nanmean(x),nanmedian(x),nanvar(x, axis=None, ddof=0),nanstd(x, axis=None, ddof=0)nanpercentile(x, q),nanquantile(x, q)— like theirnumpycounterparts, always returning floatsnanaverage(x, axis=None, weights=None, returned=False)— weighted mean excluding weights at missing positions
Selection:
nanfirst(x, axis=0),nanlast(x, axis=-1)— first/last valid value along an axis
Element-wise binary operations:
nanmaximum(x, y),nanminimum(x, y)— asnp.maximum/np.minimum, but picking the valid value wherever one input is missingnanclip(x, a_min=None, a_max=None)— asnp.clip, but preserving missing values and treating missing bounds as unbounded
Comparison:
nanequal(x, y)— element-wise equality, treating missing values as an ordinary valuenanclose(x, y, delta=sys.float_info.epsilon)— element-wise closeness with tolerancedelta
The library ships two interchangeable implementations:
intnan_np— based purely on vectorizednumpyoperationsintnan_numba— JIT-compiled with numba for functions that allow major speed gains
Both provide the identical API. On import, the numba implementation is
automatically selected whenever numba is installed and importable; otherwise
the numpy implementation is used. This makes numba an optional runtime
dependency. Compiled numba kernels are cached on disk, so no recompilation
overhead occurs after the first use.
To get the accelerated implementation, install the numba extra:
pip install "intnan[numba]"The project uses uv for dependency management:
git clone https://github.com/ml31415/intnan
cd intnan
uv sync # create virtualenv and install all dependencies
uv run pre-commit install # optional: run lint, format and type checks on every commit
uv run pytest # run the test suite, with line coverage via pytest-cov
uv run ruff check . && uv run ruff format --check . # lint and format check
uv run mypy # type checkTests are run against both implementations and a range of dtypes
(int32, int64, float32, float64) in the CI on Python 3.12 through 3.15.
BSD 3-Clause, see LICENSE.txt.