OCS/ODF: Ceph internal cluster is FULL xx full osd(s)
Environment
Red Hat OpenShift Container Platform (OCP) 4.x
Red Hat OpenShift Container Storage (OCS) 4.x
Red Hat OpenShift Data Foundation (ODF) 4.x
Red Hat Ceph Storage (RHCS) 5.x
Red Hat Ceph Storage (RHCS) 6.x
Red Hat Ceph Storage (RHCS) 7.x
Ceph (RADOS) Block Devices (RBD)
Ceph File System (CephFS)
Issue
Ceph internal cluster is "FULL", xx full osd(s)
How to delete stale / orphan cephfs clones and snapshots
Resolution
- A ReadOnly OCS or ODF cluster can be fixed by either increasing the Storage Capacity of the cluster or by deleting unwanted data from the cluster.
Increasing the OCS/ODF Cluster Capacity
The Storage Capacity can be increased: Scaling Storage for ODF.
Deleting Unwanted Data
For CephFS: How to delete stale / orphan cephfs clones and snapshots
-
RBD volumes over subscribe Ceph
- For RBD: RBD volumes over subscribe Ceph capacity, unexpected capacity growth
- This KCS also links to another KCS to find and remove orphaned RBD volumes and snapshots
- This KCS also links to Red Hat ODF Documentation to schedule fstrim on a regular basis
-
To delete any data, first the ODF/Ceph cluster must have the full ratio changed to alleviate the
FULLstate.- Even one full OSD will cause all writes and deletes to fail
- Follow KCS Article #4628891 or KCS Article #4870821 to access the Ceph CLI.
Increasing the Full and Near Full ratios
- After ODF version 4.20 for increasing ratios please refer the Documentation Setting Ceph OSD full thresholds
- For versions until 4.19 or before you can execute the commands bellow to increase the near full and full ratios. The commands will increase all ratios, remember that full ratio needs to be bigger than backfill full which needs to be bigger than near full.
NOTE: Do not increase the FULL or NEARFULL or BACKFILLFULL threshold above 88% without engaging Red Hat Support.
$ ceph osd set-full-ratio 0.88
$ ceph osd set-backfillfull-ratio 0.83
$ ceph osd set-nearfull-ratio 0.78
Fixing full error
- To get the cluster out of
FULL, delete as much unwanted data as possible:- Actually deleting data;
- Running fstrim on RBD volumes and scheduling fstrim via a cron jobs;
- Removing orphaned RBD and CephFS volumes.
- Once the data is deleted and the OCS/ODF cluster is out of
FULLstate, revert the ratios to default.
$ ceph osd set-full-ratio 0.85
$ ceph osd set-backfillfull-ratio 0.80
$ ceph osd set-nearfull-ratio 0.75
Root Cause
- The backend storage,
Red Hat Ceph, is full and a capacity should be increased. - Lack of Monitoring and Capacity Plan
- There are
orphanedCephFS Clones and/or Snapshots. - Ceph RBD volumes need
fstrimexecuted to reclaim space.
Diagnostic Steps
Check Ceph usage inside Rook Toolbox
- Use KCS Solution #4628891 to enable rook-ceph-toolbox.
Checking ceph status:
$ oc rsh `oc get pods -l app=rook-ceph-tools -o name`
sh-5.1$ bash
bash-5.1$ ceph status
- Inside rook-ceph-toolbox pod run the commands
ceph df; ceph osd df tree. Check if they are OK as the examples below:
bash-5.1$ ceph df
--- RAW STORAGE ---
CLASS SIZE AVAIL USED RAW USED %RAW USED
ssd 1000 GiB 891 GiB 109 GiB 109 GiB 10.89
TOTAL 1000 GiB 891 GiB 109 GiB 109 GiB 10.89
--- POOLS ---
POOL ID PGS STORED OBJECTS USED %USED MAX AVAIL
ocs-storagecluster-cephobjectstore.rgw.buckets.index 1 8 30 KiB 33 91 KiB 0 247 GiB
ocs-storagecluster-cephobjectstore.rgw.meta 2 8 3.7 KiB 27 160 KiB 0 247 GiB
.rgw.root 3 8 5.3 KiB 23 176 KiB 0 247 GiB
ocs-storagecluster-cephobjectstore.rgw.buckets.non-ec 4 8 0 B 0 0 B 0 247 GiB
ocs-storagecluster-cephobjectstore.rgw.otp 5 8 0 B 0 0 B 0 247 GiB
ocs-storagecluster-cephobjectstore.rgw.control 6 8 0 B 8 0 B 0 247 GiB
ocs-storagecluster-cephobjectstore.rgw.log 7 8 250 KiB 375 2.2 MiB 0 247 GiB
.mgr 8 1 299 KiB 2 904 KiB 0 247 GiB
ocs-storagecluster-cephobjectstore.rgw.buckets.data 9 32 507 MiB 328 1.5 GiB 0.20 247 GiB
ocs-storagecluster-cephfilesystem-metadata 10 16 158 MiB 99 474 MiB 0.06 247 GiB
ocs-storagecluster-cephfilesystem-data0 11 32 33 GiB 12.80k 100 GiB 11.89 247 GiB
ocs-storagecluster-cephblockpool 12 32 1.1 GiB 904 3.4 GiB 0.46 247 GiB
default.rgw.log 13 1 5.0 KiB 34 272 KiB 0 371 GiB
default.rgw.control 14 1 0 B 8 0 B 0 371 GiB
default.rgw.meta 15 1 0 B 0 0 B 0 371 GiB
rep2 16 1 19 B 2 8 KiB 0 371 GiB
.nfs 17 1 1.4 KiB 3 12 KiB 0 247 GiB
bash-5.1$ ceph osd df tree
ID CLASS WEIGHT REWEIGHT SIZE RAW USE DATA OMAP META AVAIL %USE VAR PGS STATUS TYPE NAME
-1 0.97659 - 1000 GiB 109 GiB 107 GiB 6.8 MiB 2.1 GiB 891 GiB 10.89 1.00 - root default
-4 0.48830 - 500 GiB 54 GiB 53 GiB 3.5 MiB 1.0 GiB 446 GiB 10.89 1.00 - rack rack0
-3 0.48830 - 500 GiB 54 GiB 53 GiB 3.5 MiB 1.0 GiB 446 GiB 10.89 1.00 - host odf0-libvirt2-ocpcluster-cc
1 ssd 0.48830 1.00000 500 GiB 54 GiB 53 GiB 3.5 MiB 1.0 GiB 446 GiB 10.89 1.00 174 up osd.1
-12 0.48830 - 500 GiB 54 GiB 53 GiB 3.3 MiB 1.0 GiB 446 GiB 10.89 1.00 - rack rack1
-11 0.48830 - 500 GiB 54 GiB 53 GiB 3.3 MiB 1.0 GiB 446 GiB 10.89 1.00 - host odf1-libvirt2-ocpcluster-cc
2 ssd 0.48830 1.00000 500 GiB 54 GiB 53 GiB 3.3 MiB 1.0 GiB 446 GiB 10.89 1.00 174 up osd.2
-8 0 - 0 B 0 B 0 B 0 B 0 B 0 B 0 0 - rack rack2
TOTAL 1000 GiB 109 GiB 107 GiB 6.8 MiB 2.1 GiB 891 GiB 10.89
MIN/MAX VAR: 1.00/1.00 STDDEV: 0
This solution is part of Red Hat’s fast-track publication program, providing a huge library of solutions that Red Hat engineers have created while supporting our customers. To give you the knowledge you need the instant it becomes available, these articles may be presented in a raw and unedited form.