Skip to content

Latest commit

 

History

History
572 lines (419 loc) · 15.4 KB

File metadata and controls

572 lines (419 loc) · 15.4 KB

Home › Troubleshooting Guide

Troubleshooting Guide

Solutions for common cachekit issues and error messages


Common Errors

Circuit Breaker Errors

Issue: Circuit breaker is open and calls run uncached

What it means:

  • Five failures in total since the process started (five is the default failure_threshold; successes do not reset the count): exceptions raised by the decorated function itself, cached entries that fail to deserialize or decrypt (cachekit attempts to evict each one and the call recomputes, but it still counts; with fail-closed on, an authentication failure raises, keeps the entry, and does not count), or a failure to create the backend client. Backend read and write failures do not currently count
  • Caching is disabled for this function until the process restarts

Solutions:

  1. Check Redis availability:
redis-cli ping
# Should output: PONG
  1. Verify Redis connection string:
# Check what URL is being used
env | grep REDIS
export CACHEKIT_REDIS_URL=redis://localhost:6379/0
  1. Restart the process to reset the breaker:
  • An open breaker does not currently close on its own (a known defect), so it stays open until the process restarts
  • While open, sync functions run without caching. Async functions currently raise UnboundLocalError while the breaker is open (a known defect) — see Circuit breaker open
  1. Increase timeout if network is slow (both default to 5.0 seconds):
export CACHEKIT_SOCKET_TIMEOUT=10.0
export CACHEKIT_SOCKET_CONNECT_TIMEOUT=10.0

Exceptions raised by your own function reach the caller unchanged, with one caveat: @cache treats a BackendError raised by your function as a backend failure and may call the function a second time, so the caller gets the second call's result or exception. Your function's exceptions also count toward the breaker's failure_threshold (five by default): that many in total open the breaker for that function and stop caching it, even with a healthy backend. For some async configurations only a BackendError counts.

Serialization Failures

Issue: results are never cached, and the log shows Serialization failed with ...: TypeError followed by Failed to store in backend cache for ...: SerializationError

What it means:

  • Cache attempted to serialize function result
  • Data type is not compatible with chosen serializer
  • The default serializer (MessagePack) supports None, bool, int, float, str, bytes, list, tuple, dict, datetime, date and time
  • The call still returns the result, except with @cache(interop=...), where an unsupported type raises InteropError — see Serialization unsupported type

Solutions:

  1. For custom objects, convert to a supported type before caching:
from cachekit import cache

# Convert to dict before caching
@cache
def get_custom_object():
    obj = MyCustomClass()
    return obj.__dict__  # or use obj.to_dict() / dataclasses.asdict(obj)
  1. For DataFrames, use ArrowSerializer:
from cachekit import cache
from cachekit.serializers import ArrowSerializer
import pandas as pd

@cache(serializer=ArrowSerializer())
def get_dataframe():
    return pd.DataFrame({"a": [1, 2, 3]})
  1. For JSON-compatible data, use default (MessagePack):
# Default serializer handles: dict, list, str, int, float, bool, None
@cache()
def get_json_data():
    return {"key": "value", "count": 42}
  1. For Pydantic models, convert to dict first:
from pydantic import BaseModel
from cachekit import cache

class User(BaseModel):
    id: int
    name: str

# Convert model to dict before caching
@cache()
def get_user(user_id: int) -> dict:
    user = fetch_user_model(user_id)  # Returns Pydantic model
    return user.model_dump()  # Explicit conversion

Why not auto-detect Pydantic models? See Serializer Guide - Caching Pydantic Models for the detailed rationale.

Connection Issues

Issue: Redis connection timeout or refused

@cache does not raise these: it logs the failure and runs the function uncached. The log line names a BackendError wrapping one of these redis-py errors (see Connection Errors):

Error 111 connecting to localhost:6379. Connection refused.
Timeout connecting to server
Timeout reading from ...

Solutions:

  1. Start Redis locally:
# Using Docker (recommended)
docker run -d -p 6379:6379 redis:latest

# Verify connection
redis-cli ping
# Output: PONG
  1. Verify connection URL:
import redis

# Test connection before using decorator
try:
    r = redis.from_url("redis://localhost:6379/0")
    print(r.ping())
except Exception as e:
    print(f"Connection failed: {e}")
  1. Check firewall/network:
# On same machine
redis-cli -h localhost -p 6379 ping

# Across network (replace host)
redis-cli -h redis-server.example.com -p 6379 ping
  1. For timeout issues, increase timeout values (both default to 5.0 seconds):
export CACHEKIT_SOCKET_TIMEOUT=10.0
export CACHEKIT_SOCKET_CONNECT_TIMEOUT=10.0
  1. Verify Redis is running:
# Check if port 6379 is listening
netstat -tulpn | grep 6379
# or
lsof -i :6379
Encryption Issues

Issue: Decryption failures or key-related errors

Error messages (types and when each raises: Encryption Errors):

cache.secure requires master_key parameter or CACHEKIT_MASTER_KEY environment variable
CACHEKIT_MASTER_KEY must be hex-encoded: ...
CACHEKIT_MASTER_KEY must be at least 32 bytes (256 bits). Got ... bytes. ...
Decryption failed: ...

Solutions:

See Zero-Knowledge Encryption - Troubleshooting

Common causes:

  1. Master key not set when using @cache.secure()
  2. Master key format invalid (not hex-encoded)
  3. Master key rotated (can't decrypt old cached data)
  4. Data corruption during storage/retrieval

Quick fix:

# First-time setup only: generate a valid encryption key
export CACHEKIT_MASTER_KEY=$(openssl rand -hex 32)

# Restart application
python app.py

If the key was rotated, do not generate a new key or flush: keep the old key decrypt-only in CACHEKIT_PREVIOUS_MASTER_KEYS and follow the key rotation runbook.


CachekitIO Backend Issues

Can't Connect to cachekit.io

Issue: Requests to cachekit.io fail immediately or time out

What it means:

  • API key not configured
  • Wrong endpoint URL
  • Network/firewall blocking outbound HTTPS

Solutions:

  1. Verify API key is set:
echo $CACHEKIT_API_KEY
# Should output your key — if blank, set it:
export CACHEKIT_API_KEY=your_api_key_here
  1. Check the API URL:
# Default — leave unset unless self-hosting
echo $CACHEKIT_API_URL
# Expected: unset or https://api.cachekit.io
  1. Test network connectivity:
curl -sf https://api.cachekit.io/healthz
# Should return 200 OK — if it hangs, check firewall/proxy
401 Unauthorized

Issue: cachekit.io returns 401 Unauthorized

What it means:

  • API key is invalid, revoked, or expired
  • Key is set but doesn't match the project

Solutions:

  1. Confirm the key is correct:
# Compare against the key shown in your cachekit.io dashboard
echo $CACHEKIT_API_KEY
  1. Request a new key at cachekit.io and rotate:
export CACHEKIT_API_KEY=new_key_here
  1. Check for trailing whitespace or newlines if the key was copy-pasted:
python -c "import os; k=os.getenv('CACHEKIT_API_KEY',''); print(repr(k))"
# Key must not start/end with spaces or \n
429 Rate Limited

Issue: cachekit.io returns 429 Too Many Requests

What it means:

  • Request rate exceeds your plan's limit
  • Burst traffic spike hitting per-second cap

Solutions:

  1. Reduce cache miss rate (more hits = fewer upstream calls):
# Increase TTL to reduce backend round-trips
@cache(ttl=3600)  # 1-hour TTL instead of short TTL
def expensive_query(id):
    return fetch(id)
  1. Sync calls do not retry a rate-limited request — the call runs uncached, and the failure does not count toward the circuit breaker. Async calls currently retry the lock request for about 5 seconds plus request time on a 429 (a known defect, see CachekitIO HTTP Errors). If you're hitting 429 consistently, reduce request concurrency or upgrade your plan.

  2. Check your current usage at cachekit.io dashboard.

SSRF Rejection

Issue: Request rejected with SSRF protection error

What it means:

  • A custom CACHEKIT_API_URL points to an internal/private host
  • SSRF protection blocks requests to non-allowlisted destinations
  • Only api.cachekit.io is permitted by default

Solutions:

  1. Use the default endpoint (unset any custom URL):
unset CACHEKIT_API_URL
  1. If self-hosting, confirm your host is correctly configured and reachable:
export CACHEKIT_API_URL=https://your-self-hosted-endpoint.example.com
curl -sf $CACHEKIT_API_URL/healthz
  1. Never point CACHEKIT_API_URL at localhost or internal IPs — these are blocked by SSRF protection regardless of environment.
Connection Timeout

Issue: Requests to cachekit.io hang and eventually time out

What it means:

  • High network latency between your environment and api.cachekit.io
  • Timeout configured too low for your network conditions
  • Transient outage or overloaded backend

Solutions:

  1. Check configured timeout (the CACHEKIT_SOCKET_* variables apply to Redis only):
echo $CACHEKIT_TIMEOUT
  1. Increase timeout for high-latency environments:
export CACHEKIT_TIMEOUT=10.0
  1. Measure actual latency:
curl -o /dev/null -s -w "Connect: %{time_connect}s  Total: %{time_total}s\n" \
    https://api.cachekit.io/healthz
  1. A timeout does not fail the call: the request is logged and the function runs uncached. Timeouts do not count toward the circuit breaker. Async calls currently retry the lock request for about 5 seconds plus request time before running the function (a known defect, see CachekitIO HTTP Errors).

Error Reference

The errors cachekit raises or logs, with whether each reaches your code, are in the Error Reference.


Recovery Strategies

Cache Invalidation

Clear entire cache:

redis-cli FLUSHDB

Clear by namespace (if implemented):

from cachekit import cache

@cache(namespace="users")
def get_user(user_id):
    return fetch_user(user_id)

# Manual invalidation
# Note: Current cachekit doesn't provide built-in invalidation
# Clear Redis and re-cache on next call
redis-cli FLUSHDB

Per-function cache clearing (workaround):

from cachekit import cache
import redis

r = redis.from_url("redis://localhost:6379/0")

def invalidate_user_cache(user_id):
    key = f"users:get_user:{user_id}"
    r.delete(key)

@cache(namespace="users")
def get_user(user_id):
    return fetch_user(user_id)

# Invalidate when user data changes
user = update_user(user_id, data)
invalidate_user_cache(user_id)
Graceful Degradation

Fallback when cache fails:

from cachekit import cache
import logging

logger = logging.getLogger(__name__)

@cache(ttl=3600)
def expensive_operation(x):
    try:
        return compute_expensive_result(x)
    except Exception as e:
        logger.warning(f"Computation failed: {e}")
        # Return fallback value or raise
        return fallback_value(x)

Check cache health:

import redis

def is_redis_healthy():
    try:
        r = redis.from_url("redis://localhost:6379/0")
        r.ping()
        return True
    except Exception:
        return False

# Use in monitoring
if not is_redis_healthy():
    logger.warning("Redis unavailable - cache disabled")
Health Monitoring

Monitor cache hits/misses:

from cachekit import cache
import time

cache_stats = {"hits": 0, "misses": 0}

@cache(ttl=3600)
def monitored_function(x):
    return expensive_operation(x)

# Manual tracking (built-in metrics coming soon)
def get_hit_rate():
    total = cache_stats["hits"] + cache_stats["misses"]
    if total == 0:
        return 0
    return cache_stats["hits"] / total

Health check endpoint:

import redis
from flask import jsonify

@app.route("/health/cache")
def cache_health():
    try:
        r = redis.from_url("redis://localhost:6379/0")
        r.ping()
        return jsonify({"status": "healthy"}), 200
    except Exception as e:
        return jsonify({"status": "unhealthy", "error": str(e)}), 503

Debugging

Enable detailed logging
import logging

# Set cachekit to DEBUG level
logging.getLogger("cachekit").setLevel(logging.DEBUG)

# Set Redis client to DEBUG level
logging.getLogger("redis").setLevel(logging.DEBUG)

# View logs
logging.basicConfig(level=logging.DEBUG)

With the CachekitIO backend, cachekit keeps the hpack logger (HTTP/2 header encoding) at INFO even under a DEBUG root, because its DEBUG output decodes to your API key and lock tokens. Read transport logs before you set it to DEBUG.

Check cache key format
from cachekit.core import generate_cache_key

# See what key is generated for function
key = generate_cache_key("get_user", (123,), {})
print(f"Cache key: {key}")
# Output: cachekit:get_user:abc123def456...
Test serializer independently
from cachekit.serializers import StandardSerializer

serializer = StandardSerializer()

# Test serialization
data = {"key": "value"}
encoded = serializer.serialize(data)
print(f"Encoded: {encoded[:50]}...")

# Test deserialization
decoded = serializer.deserialize(encoded)
print(f"Decoded matches: {decoded == data}")

See Also