Skip to content

Add a device-less backend - #66

Merged
davschneller merged 6 commits into
masterfrom
davschneller/no-device2
Sep 28, 2026
Merged

davschneller merged 6 commits into
masterfrom
davschneller/no-device2

Conversation

@davschneller

Copy link
Copy Markdown
Contributor

To simplify compiling for CPU and GPU.

Supersedes #48 .

Every backend instantiates the member templates of device::Algorithms
explicitly, and each of them spelled out its own list of 36 instantiations.
algorithms/Instantiations.h now holds the type lists together with one
instantiation macro per member template; the CUDA/HIP and the SYCL backend
use them. A backend that is added later thus exports the same symbols by
construction.

No functional change: compiled against Algorithms.h, the explicit
instantiations and the macros emit the same 36 symbols for both backends.

AI-generated. Model: Claude Opus 5.5
DeviceInstance::instance() returns the singleton; api() and algorithms()
return references to the API and the algorithms of the instance.

instance() is defined in device.cpp. All shared objects of a process thus
see the same instance, also those compiled with hidden visibility, as
Python modules built by pybind11 are. api() and algorithms() are const
members that still give mutable access, as a pointer member does; call
sites that hold a const reference to the instance keep working.

This changes the interface for all consumers: getInstance() becomes
instance(), api-> becomes api(). and algorithms. becomes algorithms().
The tests and the examples are migrated accordingly.

AI-generated. Model: Claude Opus 5.5
With DEVICE_BACKEND=none, the device library implements its API on the
host; code that uses the API thus builds and runs without any GPU
toolchain or GPU. DEVICE_ARCH is not needed for it.

The host backend has exactly one device, the host itself. Device memory
is host memory, aligned to 128 bytes as on the GPU backends, with the same
bookkeeping. All work runs synchronously on the calling thread: streams
and events only serve as handles, host functions run right away, and
streamWaitMemory makes the caller wait. Graph capturing is not available.
The algorithms are plain loops with the semantics of the GPU kernels.

The complete test suite passes on it, also under AddressSanitizer and
UndefinedBehaviorSanitizer, and the library only depends on the C and C++
runtime libraries.

AI-generated. Model: Claude Opus 5.5
The common test suite does not cover events, streamWaitMemory, the
order of several host functions on one stream, the alignment and the
bookkeeping of allocations, or copies with a pitch that differs from the
width. tests/host.cpp checks these for the host backend, together with
its properties and the neutral elements of the reductions; it is only
built with DEVICE_BACKEND=none.

Both a reduction that starts the maximum at the smallest positive value
and a 2D copy that ignores the pitches make these tests fail, while the
common suite still passes.

AI-generated. Model: Claude Opus 5.5
Adds DEVICE_BACKEND=none to the build matrix. The job needs no GPU; it runs
on a GitHub-hosted runner in the CPU image that SeisSol uses as well, and
it runs the complete test suite, including the tests of the host backend.

AI-generated. Model: Claude Opus 5.5
The basic example used the real type without including common.h, and it
called getMaxSharedMemSize, getMaxThreadBlockSize and checkOffloading,
none of which the API provides. It now copies initialized data to the
device, within the device and back, and fails if the result differs.

The CI job of the host backend builds and runs it.

AI-generated. Model: Claude Opus 5.5
@davschneller
davschneller merged commit 6fd5cb2 into master Sep 28, 2026
19 checks passed
@davschneller
davschneller deleted the davschneller/no-device2 branch September 28, 2026 21:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant