Skip to content

CRITICAL: All applications failing to connect to PostgreSQL databases #52

Description

@jwaldrip

Summary

All four CrystalShards applications are unable to connect to their respective PostgreSQL databases managed by CloudNativePG operator. This is a critical blocker preventing:

  • Application deployment
  • Worker functionality
  • Any database-backed features
  • Production readiness

Current Status

All applications showing database connectivity failures:

CrystalShards.org

AppDatabase: Failed to connect to database 'crystalshards_production' with username 'app'.
Check that you have access to connect to crystalshards-postgres-rw on port 5432

CrystalDocs.org

AppDatabase: Failed to connect to database 'crystaldocs_production' with username 'app'.
Check that you have access to connect to crystaldocs-postgres-rw on port 5432

CrystalGigs.com

AppDatabase: Failed to connect to database 'crystalgigs_production' with username 'app'.
Check that you have access to connect to crystalgigs-postgres-rw on port 5432

CrystalBits.org

AppDatabase: Failed to connect to database 'crystalbits_production' with username 'app'.
Check that you have access to connect to crystalbits-postgres-rw on port 5432

Evidence

Health Check Responses (2025-10-10 07:24 UTC):

Deployment Logs (Run #18399106532):

  • All apps failed health checks after 5 retry attempts
  • Each retry showed persistent database connection failures
  • Deployment correctly failed at 07:18:46 UTC

Expected Behavior

Each application should:

  1. Connect to its respective PostgreSQL cluster via CloudNativePG service
  2. Authenticate with username app and password from Kubernetes secret
  3. Access database named {app}_production
  4. Report "healthy" in health check response

Database Architecture

CloudNativePG Operator:

  • Each app has its own PostgreSQL cluster
  • Cluster service name: {app}-postgres-rw (read-write)
  • Deployed in each app's namespace
  • Managed by CloudNativePG operator

Expected Database Details:

App: crystalshards
- Namespace: crystalshards
- Service: crystalshards-postgres-rw.crystalshards.svc.cluster.local:5432
- Database: crystalshards_production
- User: app

App: crystaldocs
- Namespace: crystaldocs
- Service: crystaldocs-postgres-rw.crystaldocs.svc.cluster.local:5432
- Database: crystaldocs_production
- User: app

App: crystalgigs
- Namespace: crystalgigs
- Service: crystalgigs-postgres-rw.crystalgigs.svc.cluster.local:5432
- Database: crystalgigs_production
- User: app

App: crystalbits
- Namespace: crystalbits
- Service: crystalbits-postgres-rw.crystalbits.svc.cluster.local:5432
- Database: crystalbits_production
- User: app

Possible Root Causes

1. PostgreSQL Pods Not Running

Check:

kubectl get postgresql -A
kubectl get pods -A -l postgres-operator.crunchydata.com/cluster

Possible issues:

  • CloudNativePG operator not installed/running
  • PostgreSQL cluster CRDs not applied
  • Pod scheduling failures
  • Image pull errors

2. Database Not Initialized

Check:

kubectl get jobs -A -l app.kubernetes.io/name=db-init
kubectl logs -n crystalshards -l app.kubernetes.io/name=db-init

Possible issues:

  • db-init job never ran
  • db-init job failed
  • Database created but migrations not run
  • User/password not created

3. Incorrect Secrets

Check:

kubectl get secret -n crystalshards crystalshards-secrets
kubectl get secret -n crystalshards crystalshards-secrets -o jsonpath='{.data.database_url}' | base64 -d

Possible issues:

  • Secret not created
  • DATABASE_URL format incorrect
  • Password doesn't match database user
  • Host/port incorrect in URL

4. Network Policies Blocking

Check:

kubectl get networkpolicies -A
kubectl describe networkpolicy -n crystalshards

Possible issues:

  • Network policies blocking app → database communication
  • Service mesh blocking traffic
  • GKE firewall rules

5. CloudNativePG Operator Issues

Check:

kubectl get pods -n cnpg-system
kubectl logs -n cnpg-system -l app.kubernetes.io/name=cloudnative-pg

Possible issues:

  • Operator not installed
  • Operator pod crashlooping
  • CRD version mismatch
  • Operator RBAC issues

Required Debug Information

Need kubectl access to GKE cluster to investigate:

# 1. Check CloudNativePG operator
kubectl get pods -n cnpg-system
kubectl logs -n cnpg-system -l app.kubernetes.io/name=cloudnative-pg --tail=100

# 2. Check PostgreSQL clusters
kubectl get postgresql -A
kubectl describe postgresql -n crystalshards crystalshards-postgres

# 3. Check PostgreSQL pods
kubectl get pods -A -l postgres-operator.crunchydata.com/cluster
kubectl logs -n crystalshards -l postgres-operator.crunchydata.com/cluster=crystalshards-postgres

# 4. Check database init jobs
kubectl get jobs -A -l app.kubernetes.io/name=db-init
kubectl logs -n crystalshards -l app.kubernetes.io/name=db-init

# 5. Check secrets
kubectl get secret crystalshards-secrets -n crystalshards -o yaml

# 6. Test connectivity from app pod
kubectl exec -it -n crystalshards deployment/crystalshards-api -- sh
> apk add postgresql-client
> psql "postgresql://app:password@crystalshards-postgres-rw:5432/crystalshards_production"

# 7. Check network policies
kubectl get networkpolicies -A
kubectl describe networkpolicy -n crystalshards

Impact

HIGH - Blocks all production functionality:

Blocks Issue #51 - CrystalShards.org UI

  • UI cannot render pages that query database
  • No shard data to display
  • Browse/search pages non-functional

Blocks Issue #50 - Worker Functionality

  • Workers cannot update shard metadata
  • Workers cannot store dependencies
  • Workers cannot track build status
  • No background job processing possible

Blocks All Applications

  • No user authentication
  • No data storage or retrieval
  • No migrations can run
  • Applications effectively read-only (API-only)

Acceptance Criteria

Issue resolved when:

  • All PostgreSQL cluster pods are running
  • Database init jobs completed successfully
  • App pods can connect to databases
  • Health checks report database as "healthy"
  • Can query database from app pod
  • Migrations have run successfully
  • Deployment health checks pass

Related Issues

Priority

CRITICAL - Nothing works without databases. This is the highest priority issue blocking production readiness for all applications.

Files Referenced

Terraform:

  • /workspaces/monorepo/terraform/modules/operators/cloudnative-pg.tf
  • /workspaces/monorepo/apps/*/terraform/resource.kubernetes_secret.*.tf
  • /workspaces/monorepo/apps/*/terraform/resource.kubernetes_job.db_init.tf

Application Config:

  • /workspaces/monorepo/apps/*/config/database.cr
  • /workspaces/monorepo/apps/*/src/actions/api/health.cr

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingcriticalCritical priorityinfrastructureInfrastructure related

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions