当前位置:首页>Linux>Linux Btrfs文件系统:故障排查与恢复

Linux Btrfs文件系统:故障排查与恢复

  • 2026-10-11 07:01:00
Linux Btrfs文件系统:故障排查与恢复

字数 1777,阅读大约需 9 分钟

Btrfs 的修复能力比 ext4 强得多(COW + 事务 + checksum),但仍有各种故障场景。本篇整理诊断流程、应急修复与数据恢复。

一、ZFS 排错的"急救箱"

# 1. 看文件系统状态
$ btrfs filesystem show /mnt
$ btrfs filesystem show /mnt --all-devices

# 2. 看设备状态

$ btrfs device stats /mnt

# 3. 看挂载情况

$ mount | grep btrfs

# 4. 看 subvolume

$ btrfs subvolume list /mnt
$ btrfs subvolume show /mnt/data

# 5. 看分配

$ btrfs filesystem usage /mnt
$ btrfs filesystem df /mnt

# 6. 看错误

$ btrfs device stats /mnt

# 7. 系统日志

$ dmesg | grep -i btrfs
$ journalctl -u btrfs-*

二、故障速查表

故障 1:filesystem 变成只读

# 现象
$ touch /mnt/test
touch
: cannot touch '/mnt/test': Read-only file system

原因:

  • • 严重错误触发内核保护
  • • 设备错误无法恢复
  • • metadata 损坏

修复:

# 1. 看错误
$ dmesg | tail -50
$ btrfs device stats /mnt

# 2. 重新挂载为读写(如果错误已恢复)

$ mount -o remount,rw /mnt

# 3. 或先 readonly 模式

$ mount -o remount,ro /mnt

# 4. 备份数据(重要!)

$ rsync -av /mnt/ /backup/

# 5. 卸载重挂

$ umount /mnt
$ mount /dev/sdb /mnt

故障 2:superblock 错误

# 现象
$ mount /dev/sdb /mnt
mount: /mnt: wrong fs type, bad option, bad superblock on /dev/sdb

# 修复

$ btrfs rescue super-recover /dev/sdb
# 或

$ btrfs rescue zero-log /dev/sdb

故障 3:设备缺失

# 现象
$ btrfs filesystem show /mnt
# 显示 missing


# 修复

$ btrfs device remove missing /mnt
# 然后物理加新盘

$ btrfs device add /dev/sdb_new /mnt
$ btrfs balance start -dusage=0 /mnt

故障 4:scrub 错误

$ btrfs scrub status /mnt
...
csum_errors:             5
corrected_erasures:       0
verify_errors:           0
read_errors:             0

含义:

  • • csum_errors > 0:checksum 不一致
  • • corrected_erasures:自动修复
  • • uncorrectable_errors:无法修复

修复:

# 1. 跑 scrub
$ btrfs scrub start /mnt
$ btrfs scrub status /mnt

# 2. 多次 scrub 仍有错误:盘可能有问题

$ smartctl -a /dev/sdb

# 3. 换盘

$ btrfs replace start /dev/sdb /dev/sdb_new /mnt

故障 5:磁盘满

# 现象
$ touch /mnt/test
touch
: cannot touch '/mnt/test': No space left on device

# 看实际占用

$ btrfs filesystem df /mnt
Data, single: total=1.00GiB, used=1.00GiB

修复:

# 1. 看大文件
$ btrfs filesystem du -s /mnt/* | sort -hr | head

# 2. 删快照

$ btrfs subvolume list -s /mnt
$ btrfs subvolume delete /mnt/snapshots/old-snap

# 3. balance 整理

$ btrfs balance start -dusage=80 /mnt

# 4. 扩盘

$ btrfs device add /dev/sdX /mnt
$ btrfs balance start /mnt

故障 6:snapshot 删除失败

$ btrfs subvolume delete /mnt/snapshots/old-snap
ERROR: cannot delete '/mnt/snapshots/old-snap': Device or resource busy

修复:

# 看谁在用
$ lsof +D /mnt/snapshots/old-snap

# 退出所有使用该快照的进程


# 重新尝试

$ btrfs subvolume delete /mnt/snapshots/old-snap

故障 7:filesystem 损坏

# 现象
$ mount /dev/sdb /mnt
mount: /mnt: wrong fs type, bad option, bad superblock

# 紧急救援

$ btrfs rescue super-recover /dev/sdb
$ btrfs rescue zero-log /dev/sdb
$ btrfs rescue fix-device-size /dev/sdb

# 仍然不行

$ btrfs rescue chunk-recover /dev/sdb

故障 8:transaction abort

# dmesg
BTRFS: error (device sdb) in btrfs_commit_transaction ...

修复:

# 1. 尝试重挂
$ umount /mnt
$ mount /dev/sdb /mnt

# 2. 跑 scrub

$ btrfs scrub start /mnt

# 3. 检查 subvolume

$ btrfs subvolume list /mnt

故障 9:enospc(无空间)但 df 显示还有

# 现象
$ touch /mnt/test
ERROR: no space left on device

$ df -h /mnt
Filesystem      Size  Used Avail Use% Mounted on
/dev/sdb        10G   8G  2G   80% /mnt

原因:metadata 空间不足。df 看的是 data 空间。

修复:

# 1. 看 metadata 占用
$ btrfs filesystem df /mnt
Metadata, DUP: total=200MiB, used=200MiB
#                     ^^^^ metadata 满了


# 2. 加盘(解决根本问题)

$ btrfs device add /dev/sdc /mnt
$ btrfs balance start -mconvert=raid1 /mnt

# 3. 删大量小文件

$ btrfs filesystem du -s /mnt/some/dir
# 看哪里有大量小文件

故障 10:balance 卡住

# 看状态
$ btrfs balance status /mnt
# 显示 "paused" 或 "running" 但很久没进展

修复:

# 取消
$ btrfs balance cancel /mnt

# 重试

$ btrfs balance start -dusage=80 /mnt

# 或分批

$ for u in 90 80 70 60 50 40 30 20 10 0; do
    btrfs balance start -dusage=$u /mnt
    while
 btrfs balance status /mnt | grep -q "running"; do
        sleep
 10
    done

done

三、btrfs rescue 命令

# 查看帮助
$ btrfs rescue --help

# 主要命令

btrfs rescue super-recover <device>   # 恢复 superblock
btrfs rescue zero-log <device>        # 清空 log
btrfs rescue fix-device-size <device> # 修正 device size
btrfs rescue chunk-recover <device>   # 恢复 chunk
btrfs rescue clear-ino-cache <device> # 清空 inode cache
btrfs rescue build-dev <fsid>        # 重建 device 结构

实战 1:superblock 损坏

# 1. 看错误
$ dmesg | tail
BTRFS: failed to read superblock

# 2. 尝试 super-recover

$ btrfs rescue super-recover /dev/sdb

# 3. 还不行

$ btrfs rescue zero-log /dev/sdb
$ mount /dev/sdb /mnt

实战 2:device 异常

# 设备显示 0 字节
$ btrfs filesystem show /mnt
devid    2 size 0.00GiB used 0.00GiB path /dev/sdb

# 修复

$ btrfs rescue fix-device-size /dev/sdb
$ btrfs device scan /mnt

实战 3:磁盘大小变了(虚拟化场景)

# 虚拟机磁盘扩了,但 Btrfs 还认旧大小
$ btrfs rescue fix-device-size /dev/sdb

# 然后

$ btrfs filesystem resize max /mnt

四、错误诊断流程

文件系统报错
│
├─ mount 失败
│   ├─ "bad superblock" → btrfs rescue super-recover
│   ├─ "wrong fs type" → 检查 mkfs 是否完成
│   └─ "device busy" → 看 mount / lsof
│
├─ 读写失败
│   ├─ "Read-only file system" → 看 dmesg,备份,重挂
│   ├─ "No space left" → btrfs filesystem df 看 metadata
│   └─ "Input/output error" → scrub + dmesg
│
├─ scrub 错误
│   ├─ csum_errors > 0 → 多次 scrub
│   ├─ 持续增长 → 换盘
│   └─ super_errors → btrfs rescue
│
├─ balance 错误
│   ├─ 卡住 → 取消分批
│   └─ enospc → 扩盘
│
└─ 设备 missing
    ├─ 物理故障 → 换盘
    └─ 设备路径变 → btrfs device scan

五、灾难恢复

5.1 多个盘故障

# RAID5:1 块坏数据安全,2 块坏数据可能丢
# RAID6:2 块坏数据安全,3 块坏数据可能丢


# 操作

# 1. 先看哪个盘没坏

$ btrfs filesystem show /mnt --all-devices

# 2. 把好的盘 dd 出来(重要数据先抢救)

$ dd if=/dev/sdc of=/backup/sdc.img bs=4M

# 3. 物理换盘

# 4. 重建

$ btrfs device add /dev/sdb_new /mnt
$ btrfs balance start -dusage=0 /mnt

5.2 完全无法挂载

# 1. 试 readonly 挂载
$ mount -o ro,recovery /dev/sdb /mnt

# 2. 试 superblock 恢复

$ btrfs rescue super-recover /dev/sdb
$ mount /dev/sdb /mnt

# 3. 试清空 log

$ btrfs rescue zero-log /dev/sdb
$ mount /dev/sdb /mnt

# 4. 都失败:dd 镜像,用 photorec 恢复

$ dd if=/dev/sdb of=/backup/sdb.img bs=4M
$ apt install testdisk
$ photorec /backup/sdb.img

六、预防胜于救援

6.1 监控清单

# 1. SMART 监控(crontab)
0 3 * * * root /usr/sbin/smartctl -t short /dev/sdX

# 2. scrub 监控

0 4 1 * * root /sbin/btrfs scrub start /mnt && btrfs scrub status /mnt > /var/log/btrfs-scrub.log

# 3. 容量监控

* * * * * root /usr/local/bin/btrfs-cap-monitor.sh

# 4. 设备错误监控

* * * * * root /usr/local/bin/btrfs-error-monitor.sh

6.2 定期任务清单

任务
频率
命令
scrub
每月
btrfs scrub start /mnt
balance
加盘后
btrfs balance start /mnt
SMART 测试
每周
smartctl -t short /dev/sdX
trim (SSD)
每周
fstrim /mnt
快照清理
每月
btrfs subvolume delete

6.3 重要数据多副本

  • • Btrfs 内部:raid1 / raid5
  • • Btrfs 外部:send/recv 异地
  • • 重要业务:3-2-1 备份原则

七、本篇小结

Btrfs 排错的核心思路:

  1. 1. 不要慌,看 dmesg 和 btrfs filesystem show
  2. 2. readonly 是内核保护机制,先备份数据再操作
  3. 3. scrub 是救命的——能修就修,修不了就标记
  4. 4. btrfs rescue 工具救不了的就 dd 镜像
  5. 5. 预防 胜于救援

关键命令速记:

任务
命令
看状态
btrfs filesystem show
看错误
btrfs device stats
救 superblock
btrfs rescue super-recover
清 log
btrfs rescue zero-log
修设备大小
btrfs rescue fix-device-size
scrub
btrfs scrub start
balance
btrfs balance start
换盘
btrfs replace start

预防措施:

  • • 每月 scrub
  • • SMART 监控
  • • 异地备份(send/recv)
  • • 监控 CAP < 90%

最新文章

随机文章