Hello,
> No, it's not an unclean shutdown. Fixed. Can't it work without a journal? > Is the orphan list not a separate feature flag? In theory, yes, but architecturally, it violates the constraints of the Mach pager. Without a journal, the pager flushes dirty blocks asynchronously. If it flushes a modified orphan inode to disk before it flushes the updated list pointer (i_dtime) of the preceding node, the linked list breaks. I spent the last few days tracing this exact race condition. Fsck is much more strict when walking the orphan list than just detecting a node with 0 i_dtime. Without the strict chronological serialization barrier provided by the journal, a hard crash leaves the linked list corrupted, turning simple "zero dtime" fsck warnings into severe "corrupted block" errors and similar. Gating the orphan list behind the journal guarantees the write-ordering needed to safely maintain the on-disk pointers. > Is that really needed? We'd normally only ever call diskfs_orphan_add > when st_nlink got down to 0. Unlinked nodes often times get their blocks smashed together by ext2/pager, one unlinked file pointer starts pointing to a completely different files blocks. This is visible on unjornaled, unorphanted ext2. One can do sudo apt update && sudo apt upgrade && sudo halt And one will occasionally see interesting things in some of the unlinked files. Some pointers from one file point to other files etc. This is something i was trying to iron out for the last few days, but there is no way around it, couple of these fields need to be set in the orphan for a cleaner fsck. Basically we are racing with the pager, and some of these fields need to be altered together.. Also, linux' ext4 seems to be using orphans for truncated files too, we > might want to do that too. This is a great idea, and i think its a great next patch :) How is synchronization between the superblock, the inode, and the > journal handled? This needs to be explained in the comments. Done Can't we record that somewhere in memory? It can very quickly grow on > machine upgrade. Done, now its a doubly linked list in-memory, so add/delete are constant operations. You can now drop the XXX: you are fixing that case. Done. There are other st_nlink-- in this file Thanks, i think i got them all now. Btw the patch also reverts a "workaround" on diskfs_shutdown_pager introduced in the previous journaling patch, its not needed anymore. Thanks, Milos On Sat, Sep 19, 2026 at 10:54 AM Samuel Thibault <[email protected]> wrote: > Hello, > > Milos Nikic, le mer. 16 sept. 2026 04:47:31 -0700, a ecrit: > > diff --git a/ext2fs/ext2fs.c b/ext2fs/ext2fs.c > > index 984df0448..5b996370e 100644 > > --- a/ext2fs/ext2fs.c > > +++ b/ext2fs/ext2fs.c > > @@ -271,6 +271,8 @@ main (int argc, char **argv) > > fprintf (stderr, "ext2fs: journaling enabled on %s\n", > diskfs_disk_name); > > JRNL_LOG_DEBUG ("Global Journal Initialized at %p", > ext2_journal); > > diskfs_nput(jnode); > > + /* Recover any orphan inodes left from a previous unclean > shutdown. */ > > No, it's not an unclean shutdown. > > Anything that was still memory-mapped but unlinked when ext2fs got shut > down will be in that state. That's very common after upgrading daemons > or libraries. > > > > + ext2_recover_orphan_list (); > > > > diff --git a/ext2fs/orphan.c b/ext2fs/orphan.c > > new file mode 100644 > > index 000000000..bdd0eb90b > > --- /dev/null > > +++ b/ext2fs/orphan.c > > +/* Add inode NP to the orphan list. */ > > +void > > +diskfs_orphan_add (struct node *np) > > +{ > > + ino_t inum = np->cache_id; > > + struct ext2_inode *di; > > + diskfs_transaction_t *txn = NULL; > > + > > + if (!ext2_journal) > > + return; > > Can't it work without a journal? > > Is the orphan list not a separate feature flag? > > > + di->i_links_count = 0; > > Is that really needed? We'd normally only ever call diskfs_orphan_add > when st_nlink got down to 0. > > Also, linux' ext4 seems to be using orphans for truncated files too, we > might want to do that too. > > > + /* Update the superblock to point to this inode as the new list head. > */ > > + sblock->s_last_orphan = htole32 (inum); > > + sblock_dirty = 1; > > + > > + if (txn) > > + { > > + journal_dirty_block (txn, boffs_block (bptr_offs (di))); > > + > > + /* Sync our private sblock to the Mach disk cache so the journal > captures it */ > > + memcpy (boffs_ptr (SBLOCK_OFFS), sblock, SBLOCK_SIZE); > > + journal_dirty_block (txn, boffs_block (SBLOCK_OFFS)); > > + } > > + > > + dino_deref (di); > > + > > + diskfs_node_disknode (np)->on_orphan_list = 1; > > + pthread_mutex_unlock (&orphan_lock); > > + > > + if (txn) > > + diskfs_journal_stop_transaction (txn); > > + else > > + alloc_sync (np); > > How is synchronization between the superblock, the inode, and the > journal handled? This needs to be explained in the comments. > > > +/* Remove inode NP from the orphan list. */ > > +void > > +diskfs_orphan_del (struct node *np) > > +{ > [...] > > + else > > + { > > + /* Walk the list to find the predecessor. */ > > Can't we record that somewhere in memory? It can very quickly grow on > machine upgrade. > > > +/* Recover (clean up) the orphan list at mount time. */ > > +void > > +ext2_recover_orphan_list (void) > > +{ > > + if (diskfs_readonly) > > + { > > + ext2_warning ("orphan inodes on readonly fs; leaving for fsck"); > > + return; > > Please put this warning after checking for empty orphan list. > > > diff --git a/libdiskfs/dir-rename.c b/libdiskfs/dir-rename.c > > index 939d0b6ae..e85ef45bb 100644 > > --- a/libdiskfs/dir-rename.c > > +++ b/libdiskfs/dir-rename.c > > @@ -240,6 +240,8 @@ diskfs_S_dir_rename (struct protid *fromcred, > > diskfs_node_update (fdp, diskfs_synchronous); > > > > fnp->dn_stat.st_nlink--; > > + if (fnp->dn_stat.st_nlink == 0) > > + diskfs_orphan_add (fnp); > > fnp->dn_set_ctime = 1; > > > > diskfs_node_update (fnp, diskfs_synchronous); > > There are other st_nlink-- in this file, don't we want to orphan them? > > > diff --git a/libdiskfs/dir-renamed.c b/libdiskfs/dir-renamed.c > > index 97487ce53..8fcaffdba 100644 > > --- a/libdiskfs/dir-renamed.c > > +++ b/libdiskfs/dir-renamed.c > > @@ -212,6 +212,8 @@ diskfs_rename_dir (struct node *fdp, struct node > *fnp, const char *fromname, > > if (!err) > > { > > tnp->dn_stat.st_nlink--; > > + if (tnp->dn_stat.st_nlink == 0) > > + diskfs_orphan_add (tnp); > > tnp->dn_set_ctime = 1; > > } > > diskfs_clear_directory (tnp, tdp, tocred); > > Same here, there is another one in an error case. > > > @@ -251,6 +253,8 @@ diskfs_rename_dir (struct node *fdp, struct node > *fnp, const char *fromname, > > diskfs_dirremove (fdp, fnp, fromname, ds); > > ds = 0; > > fnp->dn_stat.st_nlink--; > > + if (fnp->dn_stat.st_nlink == 0) > > + diskfs_orphan_add (fnp); > > fnp->dn_set_ctime = 1; > > diskfs_file_update (fdp, diskfs_synchronous); > > diskfs_node_update (fnp, diskfs_synchronous); > > > > > diff --git a/libdiskfs/node-drop.c b/libdiskfs/node-drop.c > > index a12c29ad0..bb2d38e26 100644 > > --- a/libdiskfs/node-drop.c > > +++ b/libdiskfs/node-drop.c > > @@ -43,7 +43,7 @@ diskfs_drop_node (struct node *np) > > /* XXX: if the filesystem is readonly, we cannot remove the files > with no link > > You can now drop the XXX: you are fixing that case. > > > but e.g. memory mapping still in memory. This notably happens when > > upgrading packages without restarting the corresponding > processes. Fsck > > - will have to fix them. */ > > + will have to fix them or the orphan list, if implemented. */ > > if (np->dn_stat.st_nlink == 0 && !diskfs_readonly) > > { > > diskfs_check_readonly (); > > @@ -85,9 +85,13 @@ diskfs_drop_node (struct node *np) > > np->dn_stat.st_rdev = 0; > > np->dn_set_ctime = np->dn_set_atime = 1; > > diskfs_node_update (np, diskfs_synchronous); > > + diskfs_orphan_del (np); > > diskfs_free_node (np, savemode); > > } > > else > > + /* Here we don't remove the node from the orphan list > > + so that on the next restart file system has the > > + opportunity to deal with it before fsck. */ > > diskfs_node_update (np, diskfs_synchronous); > > > > fshelp_drop_transbox (&np->transbox); > > Thanks, > Samuel >
From 395ca6c1606aba7b2ceb2670487692715f070510 Mon Sep 17 00:00:00 2001 From: Milos Nikic <[email protected]> Date: Mon, 14 Sep 2026 13:57:39 -0700 Subject: [PATCH] ext2fs: Add ext3 style orphan list When the file system is in read-only mode it is not possible to remove inodes from disk. This causes issues that fsck needs to fix later. The orphan list serves as a list of nodes that have nlink == 0 in memory but didn't manage to get deleted from disk in time. Such a list gives file system an opportunity to remove them on the following restart prior to fsck run. Even if the file system itself doesn't fix it, fsck understands this implementation of the orphan list and will gracefully clean it up the first time it sees it. It works by utilizing the superblock's s_last_orphan as the head of the singly linked list on disk. If s_last_orphan is 0, the list is empty and there are no orphans. Otherwise it points to the last element added to that list. That element is the node that needs to be deleted, so by convention its i_dtime is altered to point to the next element and so on. In memory, it is represented as a doubly linked list for efficient updates. This patch also reverts changes to libdiskfs/pager.c done by the journal patch 370c37aae. (they were a workaround for the fact that orphan list wasn't yet implemented). --- ext2fs/Makefile | 2 +- ext2fs/ext2fs.c | 5 + ext2fs/ext2fs.h | 12 ++ ext2fs/ialloc.c | 3 + ext2fs/inode.c | 46 +++--- ext2fs/orphan.c | 350 ++++++++++++++++++++++++++++++++++++++++ ext2fs/pager.c | 71 +------- libdiskfs/Makefile | 2 +- libdiskfs/dir-clear.c | 4 + libdiskfs/dir-init.c | 6 + libdiskfs/dir-link.c | 4 + libdiskfs/dir-rename.c | 11 +- libdiskfs/dir-renamed.c | 13 ++ libdiskfs/dir-rmdir.c | 2 + libdiskfs/dir-unlink.c | 2 + libdiskfs/diskfs.h | 14 ++ libdiskfs/node-create.c | 1 + libdiskfs/node-drop.c | 8 +- libdiskfs/orphan.c | 40 +++++ 19 files changed, 503 insertions(+), 93 deletions(-) create mode 100644 ext2fs/orphan.c create mode 100644 libdiskfs/orphan.c diff --git a/ext2fs/Makefile b/ext2fs/Makefile index a2b0f1eef..3a8f1ada0 100644 --- a/ext2fs/Makefile +++ b/ext2fs/Makefile @@ -22,7 +22,7 @@ makemode := server target = ext2fs SRCS = balloc.c dir.c ext2fs.c getblk.c hyper.c ialloc.c \ inode.c pager.c pokel.c truncate.c storeinfo.c msg.c xinl.c \ - xattr.c journal.c + xattr.c journal.c orphan.c OBJS = $(SRCS:.c=.o) HURDLIBS = diskfs pager iohelp fshelp store ports ihash shouldbeinlibc LDLIBS = -lpthread $(and $(HAVE_LIBBZ2),-lbz2) $(and $(HAVE_LIBZ),-lz) diff --git a/ext2fs/ext2fs.c b/ext2fs/ext2fs.c index 984df0448..9bdaa2b0a 100644 --- a/ext2fs/ext2fs.c +++ b/ext2fs/ext2fs.c @@ -278,6 +278,11 @@ main (int argc, char **argv) JRNL_LOG_DEBUG ("\n[JOURNAL CHECK] No Journal flag found."); } + /* Recover unlinked but open inodes left from a previous shutdown. + It won't run unless readonly flag is false. So not for the root + filesystem and not for the unclean translator. */ + ext2_recover_orphan_list (); + /* Now that we are all set up to handle requests, and diskfs_root_node is set properly, it is safe to export our fsys control port to the outside world. */ diff --git a/ext2fs/ext2fs.h b/ext2fs/ext2fs.h index 0975457d1..5bf20e645 100644 --- a/ext2fs/ext2fs.h +++ b/ext2fs/ext2fs.h @@ -179,6 +179,13 @@ struct disknode partially allocated. */ int last_page_partially_writable; + /* True if this inode is on the ext3 orphan list (nlink=0 but still + open). The i_dtime field is used as the next pointer on disk. */ + int on_orphan_list; + /* Prev and next pointers in an in-memory doubly linked list of orphans. */ + struct node *orphan_prev; + struct node *orphan_next; + /* Index to start a directory lookup at. */ int dir_idx; }; @@ -346,6 +353,8 @@ error_t journal_dirty_block (diskfs_transaction_t * txn, block_t fs_blocknr); void journal_notify_block_changed (block_t block); + +void ext2_orphan_drop_ram_link (struct node *np); /* ---------------------------------------------------------------- */ /* Random stuff calculated from the super block. */ @@ -490,6 +499,9 @@ _dino_deref (struct ext2_inode *inode) /* Write all active disknodes into the inode pager. */ void write_all_disknodes (void); + +/* Recover (clean up) the orphan inode list at mount time. */ +void ext2_recover_orphan_list (void); /* ---------------------------------------------------------------- */ diff --git a/ext2fs/ialloc.c b/ext2fs/ialloc.c index 3cf7a2401..7be8a15b8 100644 --- a/ext2fs/ialloc.c +++ b/ext2fs/ialloc.c @@ -343,6 +343,9 @@ diskfs_alloc_node (struct node *dir, mode_t mode, struct node **node) diskfs_node_disknode (np)->info.i_next_alloc_goal = 0; diskfs_node_disknode (np)->info.i_prealloc_block = 0; diskfs_node_disknode (np)->info.i_prealloc_count = 0; + diskfs_node_disknode (np)->on_orphan_list = 0; + diskfs_node_disknode (np)->orphan_prev = NULL; + diskfs_node_disknode (np)->orphan_next = NULL; /* diskfs_node_disknode (np)->info.i_new_inode */ /* diff --git a/ext2fs/inode.c b/ext2fs/inode.c index 5fa5165e6..1b6dae06b 100644 --- a/ext2fs/inode.c +++ b/ext2fs/inode.c @@ -62,6 +62,9 @@ diskfs_user_make_node (struct node **npp, struct lookup_context *ctx) dn->dirents = 0; dn->dir_idx = 0; dn->pager = 0; + dn->on_orphan_list = 0; + dn->orphan_prev = NULL; + dn->orphan_next = NULL; pthread_rwlock_init (&dn->alloc_lock, NULL); pokel_init (&dn->indir_pokel, diskfs_disk_pager, disk_cache); @@ -74,6 +77,7 @@ diskfs_user_make_node (struct node **npp, struct lookup_context *ctx) void diskfs_node_norefs (struct node *np) { + ext2_orphan_drop_ram_link (np); if (diskfs_node_disknode (np)->dirents) free (diskfs_node_disknode (np)->dirents); assert_backtrace (!diskfs_node_disknode (np)->pager); @@ -488,27 +492,31 @@ write_node (struct node *np) info->i_flags |= EXT2_IMMUTABLE_FL; di->i_flags = htole32 (info->i_flags); - if (st->st_mode == 0) - /* Set dtime non-zero to indicate a deleted file. - We don't clear i_size, i_blocks, and i_translator in this case, - to give "undeletion" utilities a chance. */ - di->i_dtime = htole32 (di->i_mtime); - else + /* The i_dtime and other fields here are used by the orphan machinery + so we don't need to touch them here if a node is an orphan. */ + if (!diskfs_node_disknode (np)->on_orphan_list) { - di->i_dtime = htole32 (0); - di->i_size = htole32 (st->st_size); - if (sizeof (off_t) >= 8 && !S_ISDIR (st->st_mode)) - /* 64bit file size */ - di->i_size_high = htole32 (st->st_size >> 32); - di->i_blocks = htole32 (st->st_blocks); + if (st->st_mode == 0) + /* Set dtime non-zero to indicate a deleted file. */ + di->i_dtime = htole32 (di->i_mtime); + else + { + /* We don't clear i_size, i_blocks, and i_translator if mode is 0, + to give "undeletion" utilities a chance. */ + di->i_dtime = htole32 (0); + di->i_size = htole32 (st->st_size); + if (sizeof (off_t) >= 8 && !S_ISDIR (st->st_mode)) + /* 64bit file size */ + di->i_size_high = htole32 (st->st_size >> 32); + di->i_blocks = htole32 (st->st_blocks); + } + + if (S_ISCHR(st->st_mode) || S_ISBLK(st->st_mode)) + di->i_block[0] = htole32 (st->st_rdev); + else + memcpy (di->i_block, diskfs_node_disknode (np)->info.i_data, + EXT2_N_BLOCKS * sizeof di->i_block[0]); } - - if (S_ISCHR(st->st_mode) || S_ISBLK(st->st_mode)) - di->i_block[0] = htole32 (st->st_rdev); - else - memcpy (di->i_block, diskfs_node_disknode (np)->info.i_data, - EXT2_N_BLOCKS * sizeof di->i_block[0]); - diskfs_end_catch_exception (); np->dn_stat_dirty = 0; diff --git a/ext2fs/orphan.c b/ext2fs/orphan.c new file mode 100644 index 000000000..f6c277711 --- /dev/null +++ b/ext2fs/orphan.c @@ -0,0 +1,350 @@ +/* Ext3-style Orphan Inode List implementation for ext2fs. + When a file is unlinked (nlink=0) but still held open by a process, + the inode is added to the orphan list (anchored at s_last_orphan in + the superblock). Each orphaned inode uses its i_dtime field as a + "next" pointer in the singly-linked list on disk. On mount, the list is + traversed and each orphan is cleaned up (truncated and freed). + At runtime, they are organized as a doubly linked list in memory. + + Copyright (C) 2026 Free Software Foundation, Inc. + Written by Milos Nikic. + + This file is part of the GNU Hurd. + + The GNU Hurd is free software; you can redistribute it and/or modify + it under the terms of the GNU General Public License as published by + the Free Software Foundation; either version 2, or (at your option) + any later version. + + This program is distributed in the hope that it will be useful, but + WITHOUT ANY WARRANTY; without even the implied warranty of + MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU + General Public License for more details. + + You should have received a copy of the GNU General Public License + along with this program; if not, write to the Free Software + Foundation, Inc., 59 Temple Place - Suite 330, Boston, MA 02111, USA. */ + +#include "ext2fs.h" +#include "journal.h" +#include "libdiskfs/diskfs.h" +#include <pthread.h> +#include <string.h> + +/* Dedicated mutex to protect the Ext3 orphan linked list and s_last_orphan. */ +static pthread_mutex_t orphan_lock = PTHREAD_MUTEX_INITIALIZER; + +/* The in-memory head of the orphan doubly linked list. */ +static struct node *ram_orphan_head = NULL; + +/* Add inode NP to the orphan list. NP is locked by the caller. */ +void +diskfs_orphan_add (struct node *np) +{ + ino_t inum = np->cache_id; + struct ext2_inode *di; + diskfs_transaction_t *txn = NULL; + + assert_backtrace (!diskfs_readonly); + assert_backtrace (np->dn_stat.st_nlink == 0); + + if (!ext2_journal) + return; + + if (diskfs_node_disknode (np)->on_orphan_list) + return; + + /* SYNCHRONIZATION OVERVIEW: + 1. orphan_lock: Protects the in-memory doubly linked list (ram_orphan_head). + 2. global_lock: Protects the in-memory superblock modifications. + 3. Journal Transaction (txn): Guarantees that the superblock pointer and the + inode pointer hit the physical disk as a single, atomic operation. */ + txn = diskfs_journal_start_transaction (); + + pthread_mutex_lock (&orphan_lock); + + if (diskfs_node_disknode (np)->on_orphan_list) + { + pthread_mutex_unlock (&orphan_lock); + if (txn) + diskfs_journal_stop_transaction (txn); + return; + } + + /* write_node must see this before it copies info.i_data. While set: + leave i_dtime alone, and do not copy i_data, i_size, or i_blocks. + libdiskfs must call diskfs_node_update as soon as this function returns. */ + diskfs_node_disknode (np)->on_orphan_list = 1; + + ext2_debug ("adding inode %lu to orphan list", (unsigned long) inum); + + di = dino_ref (inum); + + /* Atomically link the new orphan to the head of the on-disk list. */ + pthread_spin_lock (&global_lock); + di->i_dtime = sblock->s_last_orphan; + sblock->s_last_orphan = htole32 (inum); + sblock_dirty = 1; + pthread_spin_unlock (&global_lock); + + /* Isolate the inode from standard file system deletion logic. + Zeroing the block map here prevents the Mach pager from flushing garbage or + cross-linked block pointers to the disk before the journal commits. */ + di->i_links_count = 0; + di->i_size = 0; + di->i_blocks = 0; + if (!S_ISDIR (np->dn_stat.st_mode)) + di->i_size_high = 0; + memset (di->i_block, 0, EXT2_N_BLOCKS * sizeof di->i_block[0]); + + if (txn) + { + /* Atomically bundle the superblock and the placeholder inode. + By dirtying both blocks in the same transaction, we guarantee that a + crash cannot leave a severed list chain. */ + memcpy (boffs_ptr (SBLOCK_OFFS), sblock, SBLOCK_SIZE); + journal_dirty_block (txn, boffs_block (bptr_offs (di))); + journal_dirty_block (txn, boffs_block (SBLOCK_OFFS)); + + /* Ask for a synchronous commit. */ + diskfs_journal_set_sync (txn); + } + + dino_deref (di); + + /* Maintain the in-memory doubly linked list for O(1) removals. */ + diskfs_node_disknode (np)->orphan_prev = NULL; + diskfs_node_disknode (np)->orphan_next = ram_orphan_head; + if (ram_orphan_head) + diskfs_node_disknode (ram_orphan_head)->orphan_prev = np; + ram_orphan_head = np; + + pthread_mutex_unlock (&orphan_lock); + + if (txn) + diskfs_journal_stop_transaction (txn); + else + diskfs_set_hypermetadata (0, 0); +} + +/* Remove inode NP from the orphan list. NP is locked by the caller. */ +void +diskfs_orphan_del (struct node *np) +{ + ino_t inum = np->cache_id; + diskfs_transaction_t *txn = NULL; + int update_super = 0; + + if (!ext2_journal) + return; + + if (!diskfs_node_disknode (np)->on_orphan_list) + return; + + txn = diskfs_journal_start_transaction (); + + pthread_mutex_lock (&orphan_lock); + + if (!diskfs_node_disknode (np)->on_orphan_list) + { + pthread_mutex_unlock (&orphan_lock); + if (txn) + diskfs_journal_stop_transaction (txn); + return; + } + + ext2_debug ("removing inode %lu from orphan list", (unsigned long) inum); + + struct ext2_inode *my_di = dino_ref (inum); + __u32 my_next = le32toh (my_di->i_dtime); + + /* This inode is leaving the list. i_dtime becomes a normal deletion + stamp in the caller's following write_node (mode is already 0). + We do not journal_dirty it here: that copy can run before write_node + stores the cleared block map. */ + my_di->i_dtime = 0; + dino_deref (my_di); + + struct node *prev = diskfs_node_disknode (np)->orphan_prev; + struct node *next = diskfs_node_disknode (np)->orphan_next; + + if (prev == NULL) + { + pthread_spin_lock (&global_lock); + sblock->s_last_orphan = htole32 (my_next); + sblock_dirty = 1; + pthread_spin_unlock (&global_lock); + + if (txn) + { + memcpy (boffs_ptr (SBLOCK_OFFS), sblock, SBLOCK_SIZE); + journal_dirty_block (txn, boffs_block (SBLOCK_OFFS)); + } + else + update_super = 1; + + ram_orphan_head = next; + } + else + { + struct ext2_inode *prev_di = dino_ref (prev->cache_id); + + /* prev stays on the list, so its cached i_block[] is already the + placeholder (zeros). Only i_dtime changes. */ + prev_di->i_dtime = htole32 (my_next); + if (txn) + journal_dirty_block (txn, boffs_block (bptr_offs (prev_di))); + dino_deref (prev_di); + + diskfs_node_disknode (prev)->orphan_next = next; + } + + if (next) + diskfs_node_disknode (next)->orphan_prev = prev; + + diskfs_node_disknode (np)->on_orphan_list = 0; + diskfs_node_disknode (np)->orphan_prev = NULL; + diskfs_node_disknode (np)->orphan_next = NULL; + + pthread_mutex_unlock (&orphan_lock); + + if (txn) + { + if (update_super) + diskfs_journal_set_sync (txn); + diskfs_journal_stop_transaction (txn); + } + else if (update_super) + diskfs_set_hypermetadata (0, 0); +} + +/* Recover (clean up) the orphan list at mount time. + + Only inodes with nlink == 0 are truncated and freed. Entries with + nlink != 0 (corruption, or ext3-style truncate orphans, which we do + not support yet) are just taken off the list and otherwise left alone. + + The on-disk list is untrusted: each inode number is range-checked and + the walk is bounded by the total inode count, so a cycle in the chain + cannot hang the mount. */ +void +ext2_recover_orphan_list (void) +{ + ino_t inum; + __u32 next_orphan; + __u32 steps = 0; + int count = 0; + struct ext2_inode *di; + struct node *np = NULL; + error_t err; + __u32 max_inodes = le32toh (sblock->s_inodes_count); + + next_orphan = le32toh (sblock->s_last_orphan); + + if (next_orphan == 0) + return; + + if (diskfs_readonly) + { + ext2_warning ("orphan inodes on readonly fs; leaving for fsck"); + return; + } + + /* No locks needed for the RAM list here; the filesystem is strictly + single-threaded during mount. */ + while (next_orphan != 0) + { + inum = next_orphan; + + /* Validate before touching the disk. Counting steps (not + recoveries) also bounds the nlink != 0 path below. */ + if (inum < EXT2_FIRST_INO (sblock) || inum > max_inodes + || ++steps > max_inodes) + { + ext2_warning ("corrupt or cyclic orphan list (inode %lu); " + "aborting recovery", (unsigned long) inum); + /* Break the chain so we don't re-walk garbage on every mount. */ + pthread_spin_lock (&global_lock); + sblock->s_last_orphan = 0; + sblock_dirty = 1; + pthread_spin_unlock (&global_lock); + break; + } + + /* Read the next pointer before lookup/deletion: orphan_del will + scrub this inode's i_dtime. */ + di = dino_ref (inum); + next_orphan = le32toh (di->i_dtime); + dino_deref (di); + + err = diskfs_cached_lookup (inum, &np); + if (err || !np) + { + ext2_warning ("cannot look up orphan inode %lu: %s", + (unsigned long) inum, + err ? strerror (err) : "not found"); + /* We can't safely unlink an inode we can't load. Leave the + rest of the list, including s_last_orphan, for fsck. */ + break; + } + + /* Mock the RAM list state so that diskfs_orphan_del can natively + update the disk structure and s_last_orphan. This inode is + always the current head, so prev == NULL. */ + diskfs_node_disknode (np)->on_orphan_list = 1; + diskfs_node_disknode (np)->orphan_prev = NULL; + diskfs_node_disknode (np)->orphan_next = NULL; + ram_orphan_head = np; + + if (np->dn_stat.st_nlink != 0) + { + /* Not a deleted file, so diskfs_nput would not drop it and + orphan_del would never run. Unlink it explicitly; this + scrubs i_dtime, advances s_last_orphan and clears the head. */ + ext2_warning ("orphan inode %lu has nlink > 0; unlinking from list", + (unsigned long) inum); + diskfs_orphan_del (np); + diskfs_nput (np); + continue; + } + + /* This drops the reference. Since st_nlink == 0, libdiskfs will + truncate the file and call diskfs_orphan_del (np), which + advances s_last_orphan and clears the head. */ + diskfs_nput (np); + count++; + } + + ram_orphan_head = NULL; + + if (count > 0) + { + /* Lets sync it, so that fsck doesn't try to clear the same orphans again. */ + diskfs_set_hypermetadata (1, 0); + ext2_warning ("recovered %d orphan inode(s)", count); + } +} + +void +ext2_orphan_drop_ram_link (struct node *np) +{ + pthread_mutex_lock (&orphan_lock); + + if (diskfs_node_disknode (np)->on_orphan_list) + { + struct node *prev = diskfs_node_disknode (np)->orphan_prev; + struct node *next = diskfs_node_disknode (np)->orphan_next; + + if (prev) + diskfs_node_disknode (prev)->orphan_next = next; + else + ram_orphan_head = next; + if (next) + diskfs_node_disknode (next)->orphan_prev = prev; + + diskfs_node_disknode (np)->orphan_prev = NULL; + diskfs_node_disknode (np)->orphan_next = NULL; + } + + pthread_mutex_unlock (&orphan_lock); +} diff --git a/ext2fs/pager.c b/ext2fs/pager.c index 73c7d0b3c..efc104022 100644 --- a/ext2fs/pager.c +++ b/ext2fs/pager.c @@ -964,7 +964,7 @@ pager_report_extent (struct user_pager_info *pager, void pager_clear_user_data (struct user_pager_info *upi) { - if (upi->type == FILE_DATA && upi->node) + if (upi->type == FILE_DATA) { struct pager *pager; @@ -1561,81 +1561,18 @@ diskfs_get_filemap_pager_struct (struct node *node) void diskfs_shutdown_pager (void) { - /* TODO: Implement the Ext3/Ext4 Orphan Inode List (s_last_orphan). - Currently, if a file is unlinked (nlink=0) but still held open by a - Mach pager, it will be abandoned on disk without dtime=0 if the - system halts, causing fsck to complain. This manual teardown forces - the nodes to drop synchronously before the final journal commit. - Once the Orphan List is implemented, this entire manual pager cleanup - can be safely removed. Unlinked files will be added to the superblock's - orphan list, and the OS can just pull the power. The next boot will - silently clean them up. */ - error_t shutdown_and_clear (void *v_p) + error_t shutdown_one (void *v_p) { struct pager *p = v_p; - struct user_pager_info *upi = pager_get_upi (p); - - /* First, shutdown the pager: sync and flush all dirty pages, - then destroy the port right. This must happen before we - release the node reference, because pager_sync/pager_flush - may need to access the node's allocsize and alloc_lock. */ pager_shutdown (p); - - /* After pager_shutdown, the pager has been removed from the - bucket's hash table (via ports_destroy_right). But we can - still access it because ports_bucket_iterate holds a hard - reference on our behalf. - - Now release the pager's weak node reference, mimicking what - pager_dropweak + pager_clear_user_data would do. This - ensures diskfs_drop_node runs synchronously for any unlinked - nodes before we commit the final journal transaction. - - Without this, unlinked nodes would be left in a half-deleted - state: nlink=0 on disk but dtime unset, bitmap not cleared, - and free-counts not updated — all in an uncommitted journal - transaction lost on exit(0). */ - if (upi->type == FILE_DATA && upi->node) - { - int cleared = 0; - - /* Clear the node->pager back-pointer (as pager_dropweak does) - so the assert in pager_clear_user_data is satisfied. */ - pthread_spin_lock (&node_to_page_lock); - if (diskfs_node_disknode (upi->node)->pager - && pager_get_upi (diskfs_node_disknode (upi->node)->pager) == upi) - { - diskfs_node_disknode (upi->node)->pager = NULL; - cleared = 1; - } - pthread_spin_unlock (&node_to_page_lock); - - if (cleared) - ports_port_deref_weak (p); - - /* Release the weak node reference acquired in diskfs_get_filemap. - If this is the last reference, diskfs_drop_node is called - synchronously, which sets dtime, clears the inode bitmap, - and updates free-counts. */ - diskfs_nrele_light (upi->node); - - /* Prevent pager_clear_user_data (which fires when the iterator - drops its hard ref) from double-releasing the node. */ - upi->node = NULL; - } - return 0; } - ports_bucket_iterate (file_pager_bucket, shutdown_and_clear); - - /* pager_shutdown + diskfs_nrele_light above may have triggered - diskfs_drop_node for unlinked nodes, which writes dtime, clears - the inode bitmap, updates free-counts, and starts a new journal - transaction. We MUST commit this transaction before quiescing. */ write_all_disknodes (); journal_commit_running_transaction (); + ports_bucket_iterate (file_pager_bucket, shutdown_one); + if (!ext2_journal) { error_t err = store_sync (store); diff --git a/libdiskfs/Makefile b/libdiskfs/Makefile index 2b5a4a3b9..341532466 100644 --- a/libdiskfs/Makefile +++ b/libdiskfs/Makefile @@ -52,7 +52,7 @@ OTHERSRCS = conch-fetch.c conch-set.c dir-clear.c dir-init.c dir-renamed.c \ remount.c console.c disk-pager.c \ name-cache.c direnter.c dirrewrite.c dirremove.c lookup.c dead-name.c \ validate-mode.c validate-group.c validate-author.c validate-flags.c \ - validate-rdev.c validate-owner.c priv.c get-source.c journal.c + validate-rdev.c validate-owner.c priv.c get-source.c journal.c orphan.c SRCS = $(OTHERSRCS) $(FSSRCS) $(IOSRCS) $(FSYSSRCS) $(IFSOCKSRCS) installhdrs = diskfs.h diskfs-pager.h diff --git a/libdiskfs/dir-clear.c b/libdiskfs/dir-clear.c index d61baf98d..67e3eaad9 100644 --- a/libdiskfs/dir-clear.c +++ b/libdiskfs/dir-clear.c @@ -48,6 +48,8 @@ diskfs_clear_directory (struct node *dp, /* Decrement the link count */ dp->dn_stat.st_nlink--; + if (dp->dn_stat.st_nlink == 0) + diskfs_orphan_add (dp); dp->dn_set_ctime = 1; /* Find and remove the `..' entry. */ @@ -65,6 +67,8 @@ diskfs_clear_directory (struct node *dp, /* Decrement the link count on the parent */ pdp->dn_stat.st_nlink--; + if (pdp->dn_stat.st_nlink == 0) + diskfs_orphan_add (pdp); pdp->dn_set_ctime = 1; diskfs_truncate (dp, 0); diff --git a/libdiskfs/dir-init.c b/libdiskfs/dir-init.c index 5b1fb8373..8a1f14383 100644 --- a/libdiskfs/dir-init.c +++ b/libdiskfs/dir-init.c @@ -51,6 +51,8 @@ diskfs_init_dir (struct node *dp, struct node *pdp, struct protid *cred) if (err) { dp->dn_stat.st_nlink--; + if (dp->dn_stat.st_nlink == 0) + diskfs_orphan_add (dp); dp->dn_set_ctime = 1; diskfs_node_update (dp, diskfs_synchronous); @@ -67,10 +69,14 @@ diskfs_init_dir (struct node *dp, struct node *pdp, struct protid *cred) { /* ROLLBACK '.' on Parent */ pdp->dn_stat.st_nlink--; + if (pdp->dn_stat.st_nlink == 0) + diskfs_orphan_add (pdp); pdp->dn_set_ctime = 1; diskfs_node_update (pdp, diskfs_synchronous); /* CLEANUP '.' on Child */ dp->dn_stat.st_nlink--; + if (dp->dn_stat.st_nlink == 0) + diskfs_orphan_add (dp); dp->dn_set_ctime = 1; diskfs_node_update (dp, diskfs_synchronous); return err; diff --git a/libdiskfs/dir-link.c b/libdiskfs/dir-link.c index ec3c0a3d8..608ff2183 100644 --- a/libdiskfs/dir-link.c +++ b/libdiskfs/dir-link.c @@ -121,6 +121,8 @@ diskfs_S_dir_link (struct protid *dircred, { /* Deallocate link on TNP */ tnp->dn_stat.st_nlink--; + if (tnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (tnp); tnp->dn_set_ctime = 1; diskfs_node_update (tnp, diskfs_synchronous); } @@ -132,6 +134,8 @@ diskfs_S_dir_link (struct protid *dircred, if (err) { np->dn_stat.st_nlink--; + if (np->dn_stat.st_nlink == 0) + diskfs_orphan_add (np); np->dn_set_ctime = 1; diskfs_node_update (np, diskfs_synchronous); } diff --git a/libdiskfs/dir-rename.c b/libdiskfs/dir-rename.c index 939d0b6ae..6328351ad 100644 --- a/libdiskfs/dir-rename.c +++ b/libdiskfs/dir-rename.c @@ -182,6 +182,8 @@ diskfs_S_dir_rename (struct protid *fromcred, if (!err) { tnp->dn_stat.st_nlink--; + if (tnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (tnp); tnp->dn_set_ctime = 1; diskfs_node_update (tnp, diskfs_synchronous); } @@ -197,7 +199,12 @@ diskfs_S_dir_rename (struct protid *fromcred, if (err) { if (fnp->dn_stat.st_nlink > 0) - fnp->dn_stat.st_nlink--; + { + fnp->dn_stat.st_nlink--; + if (fnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (fnp); + } + fnp->dn_set_ctime = 1; diskfs_node_update (fnp, diskfs_synchronous); pthread_mutex_unlock (&fnp->lock); @@ -240,6 +247,8 @@ diskfs_S_dir_rename (struct protid *fromcred, diskfs_node_update (fdp, diskfs_synchronous); fnp->dn_stat.st_nlink--; + if (fnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (fnp); fnp->dn_set_ctime = 1; diskfs_node_update (fnp, diskfs_synchronous); diff --git a/libdiskfs/dir-renamed.c b/libdiskfs/dir-renamed.c index 97487ce53..67174874d 100644 --- a/libdiskfs/dir-renamed.c +++ b/libdiskfs/dir-renamed.c @@ -164,6 +164,8 @@ diskfs_rename_dir (struct node *fdp, struct node *fnp, const char *fromname, { assert_backtrace (tdp->dn_stat.st_nlink > 0); tdp->dn_stat.st_nlink--; + if (tdp->dn_stat.st_nlink == 0) + diskfs_orphan_add (tdp); tdp->dn_set_ctime = 1; diskfs_node_update (tdp, diskfs_synchronous); diskfs_drop_dirstat (fnp, tmpds); @@ -177,6 +179,8 @@ diskfs_rename_dir (struct node *fdp, struct node *fnp, const char *fromname, { assert_backtrace (tdp->dn_stat.st_nlink > 0); tdp->dn_stat.st_nlink--; + if (tdp->dn_stat.st_nlink == 0) + diskfs_orphan_add (tdp); tdp->dn_set_ctime = 1; diskfs_node_update (tdp, diskfs_synchronous); @@ -184,6 +188,9 @@ diskfs_rename_dir (struct node *fdp, struct node *fnp, const char *fromname, } fdp->dn_stat.st_nlink--; + if (fdp->dn_stat.st_nlink == 0) + diskfs_orphan_add (fdp); + fdp->dn_set_ctime = 1; diskfs_node_update (fdp, diskfs_synchronous); } @@ -212,6 +219,8 @@ diskfs_rename_dir (struct node *fdp, struct node *fnp, const char *fromname, if (!err) { tnp->dn_stat.st_nlink--; + if (tnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (tnp); tnp->dn_set_ctime = 1; } diskfs_clear_directory (tnp, tdp, tocred); @@ -227,6 +236,8 @@ diskfs_rename_dir (struct node *fdp, struct node *fnp, const char *fromname, { assert_backtrace (fnp->dn_stat.st_nlink > 0); fnp->dn_stat.st_nlink--; + if (fnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (fnp); fnp->dn_set_ctime = 1; /* fnp is locked, so this is safe */ diskfs_node_update (fnp, diskfs_synchronous); @@ -251,6 +262,8 @@ diskfs_rename_dir (struct node *fdp, struct node *fnp, const char *fromname, diskfs_dirremove (fdp, fnp, fromname, ds); ds = 0; fnp->dn_stat.st_nlink--; + if (fnp->dn_stat.st_nlink == 0) + diskfs_orphan_add (fnp); fnp->dn_set_ctime = 1; diskfs_file_update (fdp, diskfs_synchronous); diskfs_node_update (fnp, diskfs_synchronous); diff --git a/libdiskfs/dir-rmdir.c b/libdiskfs/dir-rmdir.c index de288aae1..82c20a7ea 100644 --- a/libdiskfs/dir-rmdir.c +++ b/libdiskfs/dir-rmdir.c @@ -90,6 +90,8 @@ diskfs_S_dir_rmdir (struct protid *dircred, if (!error) { np->dn_stat.st_nlink--; + if (np->dn_stat.st_nlink == 0) + diskfs_orphan_add (np); np->dn_set_ctime = 1; diskfs_clear_directory (np, dnp, dircred); diskfs_file_update (np, diskfs_synchronous); diff --git a/libdiskfs/dir-unlink.c b/libdiskfs/dir-unlink.c index 4ceaec4b9..85df1c977 100644 --- a/libdiskfs/dir-unlink.c +++ b/libdiskfs/dir-unlink.c @@ -79,6 +79,8 @@ diskfs_S_dir_unlink (struct protid *dircred, np->dn_stat.st_nlink--; np->dn_set_ctime = 1; + if (np->dn_stat.st_nlink == 0) + diskfs_orphan_add (np); diskfs_node_update (np, diskfs_synchronous); if (np->dn_stat.st_nlink == 0) diff --git a/libdiskfs/diskfs.h b/libdiskfs/diskfs.h index b7f4f7896..745e7044a 100644 --- a/libdiskfs/diskfs.h +++ b/libdiskfs/diskfs.h @@ -585,6 +585,20 @@ int diskfs_journal_needs_sync (diskfs_transaction_t *txn); The default definition does nothing. */ void diskfs_journal_shutdown (void); +/* Orphan Inode List hooks. + These are called by libdiskfs when a file is unlinked (nlink drops to + 0) but still held open, and when the inode is finally freed. + Filesystems with an ext3-style Orphan List (e.g. ext2fs) should + override the weak default implementations. */ + +/* Add inode NP to the orphan list. Called when nlink drops to 0 while + the node is still held open (has hard references). NP must be locked. */ +void diskfs_orphan_add (struct node *np); + +/* Remove inode NP from the orphan list. Called when the inode is about + to be permanently freed in diskfs_drop_node. NP must be locked. */ +void diskfs_orphan_del (struct node *np); + /* The user must define this function. Sync the info in NP->dn_stat and any associated format-specific information to disk. If WAIT is true, then return only after the physicial media has been completely updated. */ diff --git a/libdiskfs/node-create.c b/libdiskfs/node-create.c index 3f30cde6d..28d2c6e34 100644 --- a/libdiskfs/node-create.c +++ b/libdiskfs/node-create.c @@ -154,6 +154,7 @@ diskfs_create_node (struct node *dir, diskfs_clear_directory (np, dir, cred); np->dn_stat.st_nlink = 0; np->dn_set_ctime = 1; + diskfs_orphan_add (np); diskfs_node_update (np, diskfs_synchronous); diskfs_nput (np); } diff --git a/libdiskfs/node-drop.c b/libdiskfs/node-drop.c index a12c29ad0..c1c6cee2f 100644 --- a/libdiskfs/node-drop.c +++ b/libdiskfs/node-drop.c @@ -40,10 +40,6 @@ diskfs_drop_node (struct node *np) mode_t savemode; diskfs_transaction_t *txn = diskfs_journal_start_transaction (); - /* XXX: if the filesystem is readonly, we cannot remove the files with no link - but e.g. memory mapping still in memory. This notably happens when - upgrading packages without restarting the corresponding processes. Fsck - will have to fix them. */ if (np->dn_stat.st_nlink == 0 && !diskfs_readonly) { diskfs_check_readonly (); @@ -84,10 +80,14 @@ diskfs_drop_node (struct node *np) np->dn_stat.st_mode = 0; np->dn_stat.st_rdev = 0; np->dn_set_ctime = np->dn_set_atime = 1; + diskfs_orphan_del (np); diskfs_node_update (np, diskfs_synchronous); diskfs_free_node (np, savemode); } else + /* Here we don't remove the node from the orphan list + so that on the next restart, the file system has the + opportunity to deal with it before fsck. */ diskfs_node_update (np, diskfs_synchronous); fshelp_drop_transbox (&np->transbox); diff --git a/libdiskfs/orphan.c b/libdiskfs/orphan.c new file mode 100644 index 000000000..fed512de3 --- /dev/null +++ b/libdiskfs/orphan.c @@ -0,0 +1,40 @@ +/* Default orphan list hooks for libdiskfs. + Provides weak no-op implementations of the orphan list functions. + Filesystems with an ext3-style orphan list (e.g. ext2fs) override these. + + Written by Milos Nikic. + Copyright (C) 2026 Free Software Foundation, Inc. + + This file is part of the GNU Hurd. + + The GNU Hurd is free software; you can redistribute it and/or modify + it under the terms of the GNU General Public License as published by + the Free Software Foundation; either version 2, or (at your option) + any later version. + + This program is distributed in the hope that it will be useful, but + WITHOUT ANY WARRANTY; without even the implied warranty of + MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU + General Public License for more details. + + You should have received a copy of the GNU General Public License + along with this program; if not, write to the Free Software + Foundation, Inc., 59 Temple Place - Suite 330, Boston, MA 02111, USA. */ + +#include "diskfs.h" + +/* Add inode NP to the orphan list. Called when nlink drops to 0 + while the node is still held open (has hard references). + NP must be locked. The default implementation does nothing. */ +void __attribute__((weak)) diskfs_orphan_add (struct node *np) +{ + /* Do nothing */ +} + +/* Remove inode NP from the orphan list. Called when the inode is + about to be permanently freed in diskfs_drop_node. + NP must be locked. The default implementation does nothing. */ +void __attribute__((weak)) diskfs_orphan_del (struct node *np) +{ + /* Do nothing */ +} -- 2.55.0
