Some checks failed
API CI / API Lint (push) Successful in 47s
API CI / API Type Check (push) Failing after 47s
API CI / API Test (push) Successful in 1m0s
API CI / API Build (push) Failing after 50s
API CI / Build Docker Image (push) Has been skipped
Build Crossplane Provider / build (push) Failing after 5m51s
CD Pipeline / Deploy to Staging (push) Failing after 29s
CI Pipeline / Lint and Type Check (push) Failing after 36s
CI Pipeline / Build (push) Has been skipped
CI Pipeline / Test Backend (push) Failing after 1m33s
CI Pipeline / Test Frontend (push) Failing after 30s
CI Pipeline / Security Scan (push) Failing after 1m16s
Crossplane Provider CI / Go Test (push) Failing after 3m23s
Crossplane Provider CI / Go Lint (push) Failing after 7m27s
Crossplane Provider CI / Go Build (push) Failing after 3m27s
Deploy to Staging / Deploy to Staging (push) Failing after 30s
Portal CI / Portal Lint (push) Failing after 21s
Portal CI / Portal Type Check (push) Failing after 21s
Portal CI / Portal Test (push) Failing after 21s
Portal CI / Portal Build (push) Failing after 22s
Test Suite / frontend-tests (push) Failing after 30s
Test Suite / api-tests (push) Failing after 49s
Test Suite / blockchain-tests (push) Failing after 30s
Type Check / type-check (map[directory:. name:root]) (push) Failing after 23s
Type Check / type-check (map[directory:api name:api]) (push) Failing after 21s
Type Check / type-check (map[directory:portal name:portal]) (push) Failing after 19s
Validate Configuration Files / validate (push) Failing after 1m52s
CD Pipeline / Deploy to Production (push) Has been skipped
Co-authored-by: Cursor <cursoragent@cursor.com>
112 lines
2.5 KiB
Markdown
112 lines
2.5 KiB
Markdown
# Root Cause Identified - Ceph Services Failed
|
|
|
|
**Date**: 2025-12-13
|
|
**Status**: ✅ **ROOT CAUSE FOUND**
|
|
|
|
---
|
|
|
|
## Key Findings
|
|
|
|
### ✅ Good News
|
|
- **Bootstrap keyring EXISTS**: `/var/lib/ceph/bootstrap-osd/ceph.keyring` ✓
|
|
- **pveceph available**: Proxmox Ceph tool is installed ✓
|
|
- **Ceph config files exist**: Both `/etc/ceph/ceph.conf` and `/etc/pve/ceph.conf` ✓
|
|
- **Ceph target active**: Systemd target is active ✓
|
|
|
|
### ❌ Critical Issues
|
|
- **Ceph monitor service FAILED**: `ceph-mon@r630-01.service` - **FAILED**
|
|
- **Ceph OSD service FAILED**: `ceph-osd@1.service` - **FAILED**
|
|
- **No monitor listening**: Port 6789 has no listeners
|
|
- **Cluster not accessible**: Timeout when trying to connect
|
|
|
|
---
|
|
|
|
## Root Cause
|
|
|
|
**Ceph services are FAILED on R630-01**. This is why:
|
|
- OSD creation hangs (can't connect to cluster)
|
|
- Commands timeout (no monitors running)
|
|
- Cannot authenticate (services not running)
|
|
|
|
---
|
|
|
|
## Solution
|
|
|
|
### Step 1: Check Service Status and Logs
|
|
|
|
```bash
|
|
# Check why services failed
|
|
systemctl status ceph-mon@r630-01.service
|
|
systemctl status ceph-osd@1.service
|
|
|
|
# Check logs
|
|
journalctl -u ceph-mon@r630-01.service -n 50
|
|
journalctl -u ceph-osd@1.service -n 50
|
|
```
|
|
|
|
### Step 2: Fix and Start Services
|
|
|
|
```bash
|
|
# Try to start monitor
|
|
systemctl start ceph-mon@r630-01.service
|
|
|
|
# Try to start OSD
|
|
systemctl start ceph-osd@1.service
|
|
|
|
# Check status
|
|
systemctl status ceph-mon@r630-01.service
|
|
systemctl status ceph-osd@1.service
|
|
```
|
|
|
|
### Step 3: Verify Cluster Accessibility
|
|
|
|
```bash
|
|
# Test cluster connectivity
|
|
ceph health
|
|
ceph mon stat
|
|
ceph osd tree
|
|
```
|
|
|
|
### Step 4: Retry OSD Creation
|
|
|
|
Once services are running:
|
|
|
|
```bash
|
|
# Option A: Use pveceph (recommended for Proxmox)
|
|
for drive in sdc sdd sde sdf sdg sdh; do
|
|
pveceph create /dev/$drive
|
|
done
|
|
|
|
# Option B: Use ceph-volume (bootstrap keyring exists)
|
|
for drive in sdc sdd sde sdf sdg sdh; do
|
|
ceph-volume lvm create --data /dev/$drive
|
|
done
|
|
```
|
|
|
|
---
|
|
|
|
## Why Services Failed
|
|
|
|
Common reasons:
|
|
1. **Configuration issues** - Corrupted config
|
|
2. **Disk issues** - OSD disk problems
|
|
3. **Cluster quorum** - Lost quorum
|
|
4. **Network issues** - Cannot reach other nodes
|
|
5. **Permission issues** - Keyring permissions
|
|
|
|
---
|
|
|
|
## Next Steps
|
|
|
|
1. **Check service logs** to understand why they failed
|
|
2. **Fix the issues** preventing services from starting
|
|
3. **Start services** and verify they stay running
|
|
4. **Test cluster connectivity**
|
|
5. **Create OSDs** on the 6x 250GB drives
|
|
|
|
---
|
|
|
|
**Last Updated**: 2025-12-13
|
|
**Status**: ✅ **ROOT CAUSE IDENTIFIED - FIX SERVICES FIRST**
|
|
|