CloudEmu can deliberately fail or slow down services in controlled, time-bounded ways — so the parts of your app that handle cloud failure can actually be exercised in tests.
This is something real cloud can't do (you can't ask AWS to fail S3 for 5 seconds) and existing emulators don't do well.
Wrap any driver with the chaos engine before handing it to the portable API or the SDK-compat HTTP server. Then declare scenarios at runtime and the chaos applies to every call that hits the wrapped driver — Go API or SDK.
Chaos is wired in-process (library mode): you wrap the driver in Go, so scenarios apply whether calls arrive through the typed Go API or an in-process SDK-compat server. This is distinct from how you integrate cloudemu with a running app — that's server mode plus an SDK endpoint override, the default for integration and E2E, which doesn't expose chaos. The httptest.NewServer below is that in-process wiring, not an instruction to spin cloudemu up in a _test.go for integration.
import (
"github.com/stackshy/cloudemu/v2"
"github.com/stackshy/cloudemu/v2/features/chaos"
"github.com/stackshy/cloudemu/v2/config"
awsserver "github.com/stackshy/cloudemu/v2/server/aws"
)
cloud := cloudemu.NewAWS()
engine := chaos.New(config.RealClock{})
defer engine.Stop()
// Wrap the S3 driver. Same wrapper works for Go API or SDK-compat path.
chaosS3 := chaos.WrapBucket(cloud.S3, engine)
srv := awsserver.New(awsserver.Drivers{S3: chaosS3})
ts := httptest.NewServer(srv)
// Apply a scenario; SDK calls during the window will fail or slow down.
engine.Apply(chaos.ServiceOutage("storage", 5*time.Second))| Scenario | What it does |
|---|---|
ServiceOutage(svc, duration) |
Every call to svc returns Unavailable until the window expires |
LatencySpike(svc, extra, duration) |
Adds extra latency on every call to svc |
ProbabilisticFailure(svc, op, err, p, duration) |
Returns err on a fraction p of calls to svc.op |
Throttle(svc, op, qps, duration) |
Returns Throttled once qps calls/sec is exceeded |
Composite(scenarios...) |
Combines several scenarios; latencies sum, first error wins |
Each call to engine.Apply returns an *Active handle with .Stop() to cancel before the natural expiry.
Chaos wrapping spans the portable service layer — 20 Wrap* helpers cover
storage, compute, database, cache, DNS, IAM, container registry, logging, event
bus, load balancer, message queue, monitoring, networking, notification,
secrets, serverless, and the ML/GenAI surfaces (SageMaker, Vertex AI, Azure AI,
Azure Search):
WrapBucket, WrapCompute, WrapDatabase, WrapCache, WrapDNS, WrapIAM,
WrapContainerRegistry, WrapLogging, WrapEventBus, WrapLoadBalancer,
WrapMessageQueue, WrapMonitoring, WrapNetworking, WrapNotification,
WrapSecrets, WrapServerless, WrapSageMaker, WrapVertexAI, WrapAzureAI,
WrapAzureSearch.
events := engine.Recorded() // every Effect that was applied
engine.Reset() // clear the buffer between test phasesSlowDegradation(latency ramps up over a window)BurstFailure(N consecutive failures)NetworkPartition(cross-service: A → B fails, B → A is fine)- Pre-built scenarios based on real cloud incidents (e.g. AWS US-East-1 2017 S3 outage)
- Cascade failures via the dependency graph