CR3、HEAD 与 Catalog:分页、Git、Iceberg 共用的「可变指针 + 不可变树」

最近在对照 Linux 分页、Git 对象模型和 Iceberg 表格式,发现它们在解同一道题:底层存储块一旦写下就不改,逻辑状态却要持续变。答案都是「一枚可变指针 + 一棵不可变树」。在这里把三套源码并排看一下。


一、当前状态只是一枚指针

共通做法是:数据本身不可变,真正会改的只有「现在指向谁」。切换可见状态,就是改这枚指针。

Linux:CR3 指向页表根

x86-64 上,CR3 存的是当前地址空间顶层页表(PGD)的物理地址,外加 PCID 等控制位。切到另一个进程的 mm_struct 时,内核把新的 next->pgd 装进 CR3:

// arch/x86/mm/tlb.c:565-582
static void load_new_mm_cr3(pgd_t *pgdir, u16 new_asid, unsigned long lam,
			    bool need_flush)
{
	unsigned long new_mm_cr3;

	if (need_flush) {
		invalidate_user_asid(new_asid);
		new_mm_cr3 = build_cr3(pgdir, new_asid, lam);
	} else {
		new_mm_cr3 = build_cr3_noflush(pgdir, new_asid, lam);
	}

	/*
	 * Caution: many callers of this function expect
	 * that load_cr3() is serializing and orders TLB
	 * fills with respect to the mm_cpumask writes.
	 */
	write_cr3(new_mm_cr3);
}

参见 arch/x86/mm/tlb.c。switch_mm_irqs_off() 在真正换 mm 时走到这里1:

// arch/x86/mm/tlb.c:947-956
	if (ns.need_flush) {
		VM_WARN_ON_ONCE(is_global_asid(ns.asid));
		this_cpu_write(cpu_tlbstate.ctxs[ns.asid].ctx_id, next->context.ctx_id);
		this_cpu_write(cpu_tlbstate.ctxs[ns.asid].tlb_gen, next_tlb_gen);
		load_new_mm_cr3(next->pgd, ns.asid, new_lam, true);

可以看到,进程切换并不搬页表,也不拷物理页。它改的是 CPU 手里那枚「当前页表根」指针。write_cr3 还是一条串行化指令,后面的 TLB 填充会按新根走。CR3 里放什么、CPU 怎么用它走路,写在 Intel SDM Volume 3A 第 4 章2;内核按这个格式填表,再 mov 进寄存器。

Git:HEAD 先指分支,分支再指 commit

.git/HEAD 通常不是 commit hash,而是一条符号引用,例如 ref: refs/heads/main。真正存 hash 的是分支文件。detach 时 HEAD 才直接写 object id。git commit 先写出新的 commit 对象,再把当前分支指针推过去:

// builtin/commit.c:1938-1946
	if (commit_tree_extended(sb.buf, sb.len, &the_repository->index->cache_tree->oid,
				 parents, &oid, author_ident.buf, NULL,
				 sign_commit, extra)) {
		rollback_index_files();
		die(_("failed to write commit object"));
	}

	if (update_head_with_reflog(current_head, &oid, reflog_msg, &sb,
				    &err)) {

参见 builtin/commit.c。git-update-ref 把这件事说得很干净:给定新旧 oid,验证旧值后再写入新值3。files backend 先把 hex oid 写进 lockfile,再 rename 成正式 ref:

// refs/files-backend.c:2059-2077
static enum ref_transaction_error write_ref_to_lockfile(struct files_ref_store *refs,
							struct ref_lock *lock,
							const struct object_id *oid,
							struct strbuf *err)
{
	static char term = '\n';
	int fd;

	fd = get_lock_file_fd(&lock->lk);
	if (write_in_full(fd, oid_to_hex(oid), refs->base.repo->hash_algo->hexsz) < 0 ||
	    write_in_full(fd, &term, 1) < 0 ||
	    fsync_component(FSYNC_COMPONENT_REFERENCE, get_lock_file_fd(&lock->lk)) < 0 ||
	    close_ref_gently(lock) < 0) {
// refs/files-backend.c:1864-1892
static int commit_ref(struct ref_lock *lock)
{
	char *path = get_locked_file_path(&lock->lk);
	...
	if (commit_lock_file(&lock->lk))
		return -1;
	return 0;
}

如上所示,读者看见的「当前分支」,始终是那枚 ref 指针;对象库里的 commit/tree/blob 写完就不会改。

Iceberg:Catalog 存 metadata 路径

Iceberg 规范把表状态放在 metadata 文件里,每次变更都写一份新文件,再用原子交换替换旧指针4:

All changes to table state create a new metadata file and replace the old metadata with an atomic swap.

指针落在哪,取决于 Catalog 实现。

Hadoop 路径表没有外部 metastore。原子提交是把临时 metadata rename 成下一版本号文件(vN.metadata.json)。源码把这句话写在注释里:

// core/src/main/java/org/apache/iceberg/hadoop/HadoopTableOperations.java:157-167
    int nextVersion = (current.first() != null ? current.first() : 0) + 1;
    Path finalMetadataFile = metadataFilePath(nextVersion, codec);
    FileSystem fs = getFileSystem(tempMetadataFile, conf);

    // this rename operation is the atomic commit operation
    renameToFinal(fs, tempMetadataFile, finalMetadataFile, nextVersion);

    LOG.info("Committed a new metadata file {}", finalMetadataFile);

    // update the best-effort version pointer
    writeVersionHint(nextVersion);

参见 HadoopTableOperations.java。version-hint.text 只是 best-effort 加速查找,丢了可以扫目录恢复;真正互斥的是「vN 这份文件是否已经存在」。renameToFinal() 发现目标已在,就当并发提交失败5。

Hive / JDBC 这类 metastore Catalog,指针是表属性里的 metadata_location:

// core/src/main/java/org/apache/iceberg/BaseMetastoreTableOperations.java:46-50
  public static final String TABLE_TYPE_PROP = "table_type";
  public static final String ICEBERG_TABLE_TYPE_VALUE = "iceberg";
  public static final String METADATA_LOCATION_PROP = "metadata_location";
  public static final String METADATA_HASH_PROP = "metadata_hash";
  public static final String PREVIOUS_METADATA_LOCATION_PROP = "previous_metadata_location";

Hive 提交前先核对「我看到的 base 路径」是不是 HMS 里当前那条,对不上就拒掉:

// hive-metastore/.../HiveTableOperations.java:304-310
      String metadataLocation = tbl.getParameters().get(METADATA_LOCATION_PROP);
      String baseMetadataLocation = base != null ? base.metadataFileLocation() : null;
      if (!Objects.equals(baseMetadataLocation, metadataLocation)) {
        throw new CommitFailedException(
            "Cannot commit: Base metadata location '%s' is not same as the current table metadata location '%s' for %s.%s",
            baseMetadataLocation, metadataLocation, database, tableName);

参见 HiveTableOperations.java。通过之后再把 metadata_location 写成新文件路径6。读者刷新 Catalog,拿到的就是新的 metadata.json。

三套系统并排看:

flowchart LR
    subgraph Linux["Linux x86"]
        CR3["CR3"] --> PGD["mm.pgd / 页表根"]
    end

    subgraph GitBox["Git"]
        HEAD["HEAD"] --> Branch["refs/heads/*"]
        Branch --> Commit["commit object"]
    end

    subgraph IcebergBox["Iceberg"]
        Cat["Catalog / vN.metadata.json"] --> Meta["metadata.json"]
        Meta --> Snap["current-snapshot-id"]
    end

    style CR3 fill:#87CEEB,stroke:#333,stroke-width:2px
    style HEAD fill:#87CEEB,stroke:#333,stroke-width:2px
    style Cat fill:#87CEEB,stroke:#333,stroke-width:2px

二、多层漏斗:按图索骥,不扫整片海

指针只解决「当前是哪一棵树」。树本身还得能快速缩小范围,否则每次访问都要遍历全部底层小块。这里的复杂度差就是 O(log N) / 分区裁剪 vs O(N) 全扫。

页表:PGD → P4D → PUD → PMD → PTE

树的形状是硬件定的。Intel SDM Volume 3A 第 4 章写分页2;Linux 按那套格式在内存里建表。CPU 查表时不跑内核。

四级还是五级,看四枚控制位2:

A logical processor uses 4-level paging if CR0.PG = 1, CR4.PAE = 1, IA32_EFER.LME = 1, and CR4.LA57 = 0. 4-level paging translates 48-bit linear addresses to 52-bit physical addresses.

A logical processor uses 5-level paging if CR0.PG = 1, CR4.PAE = 1, IA32_EFER.LME = 1, and CR4.LA57 = 1. 5-level paging translates 57-bit linear addresses to 52-bit physical addresses.

走表是一次迭代。第一张表的物理地址在 CR3 里;线性地址每次切若干位选一项,项要么指向下一张表,要么指向页框2:

The first paging structure used for any translation is located at the physical address in CR3.

With 4-level paging, each paging structure comprises 512 entries and translation uses 9 bits at a time from a 48-bit linear address. Bits 47:39 identify the first paging-structure entry, bits 38:30 identify a second, bits 29:21 a third, and bits 20:12 identify a fourth.

5-level paging is similar to 4-level paging except that 5-level paging translates 57-bit linear addresses. Bits 56:48 identify the first paging-structure entry, while the remaining bits are used as with 4-level paging.

Linux 把头文件写成同一组位移。每一级 512 项;P4D_SHIFT 39、PUD_SHIFT 30、PMD_SHIFT 21,就是上面那几组 9 位:

// arch/x86/include/asm/pgtable_64_types.h:50-80
#define PGDIR_SHIFT	pgdir_shift
#define PTRS_PER_PGD	512

#define P4D_SHIFT		39
#define MAX_PTRS_PER_P4D	512
#define PTRS_PER_P4D		ptrs_per_p4d

#define PUD_SHIFT	30
#define PTRS_PER_PUD	512

#define PMD_SHIFT	21
#define PTRS_PER_PMD	512

#define PTRS_PER_PTE	512

参见 pgtable_64_types.h。选哪一项,就是右移再掩码:

// include/linux/pgtable.h:69-72
#define pgd_index(a)  (((a) >> PGDIR_SHIFT) & (PTRS_PER_PGD - 1))

pgdir_shift 默认 39,顶层对应手册的 PML4(线性地址 bits 47:39)。打开五级后改成 48,顶层变成 PML5(bits 56:48)7:

// arch/x86/boot/compressed/pgtable_64.c:16-17,121-128
unsigned int __section(".data") pgdir_shift = 39;
...
	if (!cmdline_find_option_bool("no5lvl") &&
	    native_cpuid_eax(0) >= 7 && (native_cpuid_ecx(7) & BIT(16))) {
		l5_required = true;

		__pgtable_l5_enabled = 1;
		pgdir_shift = 48;
		ptrs_per_p4d = 512;
	}

如上所示,CPUID.(EAX=07H,ECX=0):ECX.LA57 那一位,手册 §4.1.4 和压缩内核里查的是同一位。内核文档把五级写成「在现有页表上再加一层」7。

CR3 的用法手册单独写了一节8:

Ordinary 4-level paging and 5-level paging each translate linear addresses using a hierarchy of in-memory paging structures located using the contents of CR3, which is used to locate the first paging structure. For 4-level paging, this is the PML4 table, and for 5-level paging it is the PML5 table.

Table 4-12:CR4.PCIDE = 0 时,bits 12 及以上是 4K 对齐的 PML4/PML5 物理地址。Linux 切 mm 时 load_new_mm_cr3() 算出新值,最后就是一条 mov 进 CR3:

// arch/x86/include/asm/special_insns.h:41-44
static __always_inline void native_write_cr3(unsigned long val)
{
	asm volatile("mov %0,%%cr3": : "r" (val) : "memory");
}

参见 special_insns.h。

项不存在、或保留位置了 1,硬件不翻译,直接 #PF(异常 14)9:

Accesses using linear addresses may cause page-fault exceptions (#PF; exception 14). An access to a linear address may cause a page-fault exception for either of two reasons: (1) there is no translation for the linear address; or (2) there is a translation for the linear address, but its access rights do not permit the access.

there is no translation for a linear address if the translation process for that address would use a paging-structure entry in which the P flag (bit 0) is 0 or one that sets a reserved bit.

出错线性地址在 CR2。Linux 从这条向量进来,读 CR2,再决定建表、杀进程还是修 PTE:

// arch/x86/mm/fault.c:1483-1488
DEFINE_IDTENTRY_RAW_ERRORCODE(exc_page_fault)
{
	irqentry_state_t state;
	unsigned long address;

	address = cpu_feature_enabled(X86_FEATURE_FRED) ? fred_event_data(regs) : read_cr2();
// arch/x86/include/asm/trap_pf.h:7-25
 *   bit 0 ==	 0: no page found	1: protection fault
 *   bit 1 ==	 0: read access		1: write access
 *   bit 2 ==	 0: kernel-mode access	1: user-mode access
 *   bit 3 ==				1: use of reserved bit detected
...
	X86_PF_PROT	=		BIT(0),
	X86_PF_WRITE	=		BIT(1),
	X86_PF_USER	=		BIT(2),
	X86_PF_RSVD	=		BIT(3),

可以看到,这些位就是手册 Figure 4-12 的 P / W/R / U/S / RSVD。用户地址落到 do_user_addr_fault(),再 handle_mm_fault()。内核这才按树往下走,缺哪一级就分配哪一级:

// mm/memory.c:6466-6549
	pgd = pgd_offset(mm, address);
	p4d = p4d_alloc(mm, pgd, address);
	if (!p4d)
		return VM_FAULT_OOM;

	vmf.pud = pud_alloc(mm, p4d, address);
	...
	vmf.pmd = pmd_alloc(mm, vmf.pud, address);
	...
fallback:
	return handle_pte_fault(&vmf);

参见 mm/memory.c。已知地址已经落到 PMD 时,还有一条折叠助手,把四级偏移写成一行10:

// include/linux/pgtable.h:165-168
static inline pmd_t *pmd_off(struct mm_struct *mm, unsigned long va)
{
	return pmd_offset(pud_offset(p4d_offset(pgd_offset(mm, va), va), va), va);
}

Volume 3C 不重写这套走表。它把同一组寄存器收进 VMCS 的 guest-state,进 guest 时装上去11:

VM entries load processor state from these fields and VM exits store processor state into these fields.

Control registers CR0, CR3, and CR4 (64 bits each; 32 bits on processors that do not support Intel 64 architecture).

guest 里的线性地址翻译,还是 3A 第 4 章。3A 自己点到了虚拟化:VMX 可以一次装入 CR0、CR4、IA32_EFER,切到普通 MOV / WRMSR 做不到的分页模式;HLAT 用 VMCS 里的 HLATP 代替 CR3 当第一张表12。3C 叠的是 VMM 对 CR3 的保存/恢复,不是另一套用户页表格式。

Version Control with Git 里的 blob / tree / commit、Apache Iceberg: The Definitive Guide 里的 snapshot / manifest,是软件自己定的树1314。页表这棵树,格式先写在 SDM 里;Linux 的 pgd_t / pte_t 是按那份格式填的格子。没有 Linux,别的 OS 也得按同一套 CR3 跟 CPU 打交道。MMU 不会扫进程的全部物理页,只按虚拟地址切出每一级索引。

Git:commit → tree → blob

对象类型就三种常用的(再加上 tag)15:

// object.h:99-104
enum object_type {
	OBJ_BAD = -1,
	OBJ_NONE = 0,
	OBJ_COMMIT = 1,
	OBJ_TREE = 2,
	OBJ_BLOB = 3,
	OBJ_TAG = 4,

commit 只记住一棵根 tree,以及 parent 链:

// commit.h:27-39
struct commit {
	struct object object;
	timestamp_t date;
	struct commit_list *parents;
	struct tree *maybe_tree;
	unsigned int index;
};

tree 是目录项列表,项指向下一层 tree 或 blob。git commit 用 index 上的 cache_tree 当根 tree oid,不会为没改过的子树重写对象。查一个文件是「沿路径走目录项」,不是枚举整个对象库。

Iceberg:metadata → snapshot → manifest list → manifest → Parquet

规范里的快照结构是4:

  1. metadata.json 记下 schema、partition spec,以及 current-snapshot-id
  2. 每个 snapshot 有一份 manifest list
  3. manifest list 里是若干 manifest,带分区统计和文件计数
  4. manifest 里才是 data file / delete file 路径和列度量
  5. 数据文件本身通常是 Parquet(也可以是 Avro / ORC)

Parquet 文件列表不在 metadata.json 里。规范把 manifest 写成一份不可变 Avro:列出 data file 或 delete file,以及分区、度量和跟踪信息;一个 snapshot 的这些 manifest 再由 manifest list 索引一层16:

Catalog
    └─ metadata.json          当前 snapshot id,不列数据文件
        └─ snapshot
            └─ manifest list  列出有哪些 manifest(分区摘要、文件计数)
                └─ manifest   ← data file / delete file 清单在这里
                    └─ /path/to/data-a.parquet

metadata.json 只回答「当前是哪次 snapshot」,以及这次 snapshot 的 manifest list 路径。manifest list 回答「哪些 manifest 值得打开」。真正的 file_path 在 manifest 的 data_file 里。一份 snapshot 通常有多份 manifest,每份只覆盖一部分文件。

Java 里当前快照就是按 id 取:

// core/src/main/java/org/apache/iceberg/TableMetadata.java:536-538
  public Snapshot currentSnapshot() {
    return snapshotsById.get(currentSnapshotId);
  }

扫描规划明确写了可以跳过整份 manifest17:

Manifests that contain no matching files, determined using either file counts or partition summaries, may be skipped.

实现上,ManifestEvaluator 用 manifest 的分区摘要判断这份文件里有没有可能命中的分区;过了这一关,再用 InclusiveMetricsEvaluator 看单个 data file 的列上下界。eval 返回 false,这份文件就可以不读。

叶子上的 Parquet 自己又是一层漏斗。文件尾部是 FileMetaData,里面是 row group 列表;读者先读 footer,再只打开关心的 column chunk / page18:

4-byte magic number "PAR1"
<Column chunks / row groups>
File Metadata
4-byte length of file metadata
4-byte magic number "PAR1"
struct FileMetaData {
  1: required i32 version
  2: required list<SchemaElement> schema
  3: required i64 num_rows
  4: required list<RowGroup> row_groups
  ...
}

参见 parquet.thrift。Iceberg 决定读哪些文件,Parquet footer 再决定读文件里的哪些列、哪些 row group。

三棵树并排:

flowchart TB
    subgraph PT["页表"]
        VA["虚拟地址"] --> PGD2["PGD"]
        PGD2 --> P4D["P4D"]
        P4D --> PUD["PUD"]
        PUD --> PMD["PMD"]
        PMD --> PTE["PTE / 物理页"]
    end

    subgraph GT["Git"]
        BR["branch / HEAD"] --> CM["commit"]
        CM --> TR["tree"]
        TR --> TR2["tree"]
        TR --> BL["blob"]
        TR2 --> BL2["blob"]
    end

    subgraph IB["Iceberg + Parquet"]
        CAT["Catalog"] --> MD["metadata.json"]
        MD --> SN["snapshot"]
        SN --> ML["manifest list"]
        ML --> MF["manifest"]
        MF --> PQ["Parquet file"]
        PQ --> FT["footer / row group / page"]
    end

叶子这一层,Git 的 blob 对上 Iceberg 的 data file。data file 常常是 Parquet,但表格式并不绑定它。Loeliger / Ponuthorai 那本书把 blob 写成不透明字节,连文件名都不在对象里;路径在 tree 上13。Iceberg 那本书把叶子放在 data layer:manifest 跟踪文件,行数据在 data file 里,Parquet 只是最常见的一种封装14。

下面用两边都能核对的样例走一遍。

Git 书第 2 章这份 12 字节加换行,hash 是固定的13:

$ echo "hello world" > hello.txt
$ git add hello.txt
$ echo "hello world" | git hash-object --stdin
3b18e512dba79e4c8300dd08aeb37f8e728b8dad
$ git cat-file -p 3b18e512dba79e4c8300dd08aeb37f8e728b8dad
hello world

对象落在 .git/objects/3b/18e512dba79e4c8300dd08aeb37f8e728b8dad。git add 只把内容和 pathname 记进 index,还没有 tree:

$ git ls-files -s
100644 3b18e512dba79e4c8300dd08aeb37f8e728b8dad 0 hello.txt

git write-tree 才把「名字 → blob」冻成 tree13:

$ git write-tree
68aba62e560c0ebc3396e8ae9335232cd93a3f60
$ git cat-file -p 68aba6
100644 blob 3b18e512dba79e4c8300dd08aeb37f8e728b8dad hello.txt

如上所示,hello.txt 这个名字在 tree 里,不在 blob 里。再 commit 一次,commit 对象只多记 tree、作者和时间;书里的例子是13:

tree 492413269336d21fac079d4a4672e55d5d2147ac
author Jon Loeliger <jdl@example.com> 1656932750 +0200
committer Jon Loeliger <jdl@example.com> 1656932750 +0200

Commit a file that says hello

commit hash 会因作者和时间而变,tree 可以原样复用。HEAD 通常还是符号引用:

ref: refs/heads/main

refs/heads/main 里才是那串 commit oid。

Iceberg 官方测试夹具 TableMetadataV2Valid.json 把同一角色写成 JSON。Catalog 指向这份 metadata;当前可见状态是 current-snapshot-id19:

{
  "format-version": 2,
  "table-uuid": "9c12d441-03fe-4693-9a96-a0705ddf69c1",
  "location": "s3://bucket/test/location",
  "current-snapshot-id": 3055729675574597004,
  "snapshots": [
    {
      "snapshot-id": 3051729675574597004,
      "timestamp-ms": 1515100955770,
      "sequence-number": 0,
      "summary": { "operation": "append" },
      "manifest-list": "s3://a/b/1.avro"
    },
    {
      "snapshot-id": 3055729675574597004,
      "parent-snapshot-id": 3051729675574597004,
      "timestamp-ms": 1555100955770,
      "sequence-number": 1,
      "summary": { "operation": "append" },
      "manifest-list": "s3://a/b/2.avro",
      "schema-id": 1
    }
  ]
}

parent-snapshot-id 对上 Git commit 的 parent。manifest-list 对上 commit 里的 tree:下一层索引的位置,不是行数据本身。

再往下一层,测试代码用 DataFiles.builder 造叶子。路径、大小、行数写在 builder 上,不按内容算 hash20:

// core/src/test/java/org/apache/iceberg/util/TestReachableFileUtil.java:58-63
  private static final DataFile FILE_A =
      DataFiles.builder(SPEC)
          .withPath("/path/to/data-a.parquet")
          .withFileSizeInBytes(10)
          .withRecordCount(1)
          .build();

规范把这条记录展开成 manifest 里的 data_file struct。字段 id 是固定的:100 file_path、101 file_format、103 record_count、125 lower_bounds21。写成 JSON 看结构,就是:

{
  "status": 1,
  "snapshot_id": 3055729675574597004,
  "data_file": {
    "content": 0,
    "file_path": "/path/to/data-a.parquet",
    "file_format": "PARQUET",
    "record_count": 1,
    "file_size_in_bytes": 10
  }
}

status = 1 是 ADDED。真正的 hello world 行在 Parquet 文件里,不在这条元数据里。读者要读内容,得按 file_path 打开文件,再读 footer 里的 row group。

同一份「hello world」并排看:

角色 Git 样例 Iceberg 样例
可变指针 HEAD → ref: refs/heads/main Catalog 的 metadata_location → 上面这份 metadata.json
当前根 refs/heads/main 里的 commit oid "current-snapshot-id": 3055729675574597004
目录 / 文件列表 tree 68aba62e…:100644 blob 3b18e5… hello.txt snapshot 的 manifest-list: s3://a/b/2.avro,再进 manifest
叶子怎么命名 内容 SHA:3b18e512dba79e4c8300dd08aeb37f8e728b8dad 路径:/path/to/data-a.parquet
叶子里有没有名字 blob 只有 hello world\n data file 的名字在 file_path;Parquet 里是列和行
拷一份同内容 第二个文件名仍指向同一个 blob 另一条路径就是另一个 data file,即使字节相同

可以看到,说「blob 对应 Parquet」只在「不可变载荷叶子」这层成立。更贴的说法是 blob ↔ data file,tree ↔ manifest。Git 用内容做主键,所以两个路径可以共享一个 blob;Iceberg 用路径做主键,剪枝靠 manifest 上的分区和列上下界,不靠「这份 Parquet 的 SHA 跟另一份一样」。


三、先写影子,再原子切换

修改不能直接打在别人正在读的那份数据上。三套系统都是:在旁边做好新版本,最后一步才让指针看见它。

缺页分配物理页,和写时复制(COW)不是同一条路径。前者是「这页还没有」;后者是「这页有,但是只读共享,写就要拷一份」。内核把它们分成两条 fault 路径。

Linux:do_wp_page 拷页,再换 PTE

私有映射上的写过错到 do_wp_page()。能复用就复用;必须拷的时候走 wp_page_copy()22:

// mm/memory.c:4291-4320
	if (folio && folio_test_anon(folio) &&
	    (PageAnonExclusive(vmf->page) || wp_can_reuse_anon_folio(folio, vma))) {
		...
		wp_page_reuse(vmf, folio);
		return 0;
	}
	...
	return wp_page_copy(vmf);

拷完之后,先清旧 PTE 并冲 TLB,再挂上新页。注释写得很清楚:必须先切换页表项,才能把旧页的 mapcount 减掉,否则别的进程可能在窗口里写进旧页23:

// mm/memory.c:3918-3929
		ptep_clear_flush(vma, vmf->address, vmf->pte);
		folio_add_new_anon_rmap(new_folio, vma, vmf->address, RMAP_EXCLUSIVE);
		folio_add_lru_vma(new_folio, vma);
		BUG_ON(unshare && pte_write(entry));
		set_pte_at(mm, vmf->address, vmf->pte, entry);

对这个进程来说,写操作成功了;共享这份旧页的其他进程,页表还指着原来的只读页。可见性切换发生在这一条 PTE,不是整棵页表重写。

fork 之后父子共享只读页、一方先写再触发上面这条路径,就是教科书里的进程级 COW。进程切换本身仍然只写 CR3。

Git:对象先落盘,ref 后移动

commit_tree_extended() 先拼 commit 缓冲区,校验 tree 类型,再写入对象库:

// commit.c:1729-1760
int commit_tree_extended(const char *msg, size_t msg_len,
			 const struct object_id *tree,
			 const struct commit_list *parents, struct object_id *ret,
			 ...)
{
	...
	odb_assert_oid_type(the_repository->objects, tree, OBJ_TREE);
	...
	write_commit_tree(&buffer, msg, msg_len, tree, parent_buf, nparents, author, committer, extra);

参见 commit.c。这一步失败,HEAD 不动。写成功之后,update_head_with_reflog() 才去锁 ref、写 lockfile、commit_lock_file() rename。并发更新用「期望的旧 oid」做 CAS,对不上就失败——和 Iceberg 核对 metadata_location 是同一类约束。

没改过的 blob / 子 tree 继续被新 tree 引用,这就是 Git 的 COW:只为变化路径分配新对象。

Iceberg:数据文件不可变,提交只换 metadata 指针

规范要求:文件写下去就不改;表不需要随机写。Hadoop 表才依赖 rename 实现 metadata 提交4。一次 append 大致是:

  1. 写出新的 Parquet(以及需要的 delete file)
  2. 写出新的 manifest;旧 snapshot 里还能用的 manifest 直接复用
  3. 写出新的 manifest list 和 metadata.json
  4. Catalog 原子替换指针

SnapshotProducer.commit() 先 apply() 得到新 snapshot,再 taskOps.commit(base, updated);撞上 CommitFailedException 就按 commit.num-retries 重试24:

// core/src/main/java/org/apache/iceberg/SnapshotProducer.java:480-522
  public void commit() {
    AtomicLong newSnapshotId = new AtomicLong(-1L);
    ...
                taskOps -> {
                  Snapshot newSnapshot = apply();
                  newSnapshotId.set(newSnapshot.snapshotId());
                  TableMetadata.Builder update = TableMetadata.buildFrom(base);
                  ...
                  TableMetadata updated = update.build();
                  if (updated.changes().isEmpty()) {
                    return;
                  }
                  taskOps.commit(base, updated.withUUID());
                });

重试时序列号会重分,但新 manifest 可以复用——规范把这件事设计进了「从 manifest list 继承 sequence number」4。读者在指针切换前一直看着旧 snapshot,不会看见半成品文件。

sequenceDiagram
    participant W as Writer
    participant Shadow as 影子文件
    participant Ptr as 指针CR3或HEAD或Catalog
    participant R as 并发读者

    R->>Ptr: 读当前指针
    Ptr-->>R: 旧根
    W->>Shadow: 写新页或新对象或新metadata
    Note over W,Shadow: 读者仍走旧根
    W->>Ptr: 原子切换
    W-->>R: 下次刷新才看到新根

四、历史还在,回收另做

指针往前走以后,旧树不必立刻消失。历史查询靠「还有没有人引用」;物理回收是另一次显式动作。

Linux:进程退出拆掉页表

exit_mmap() 在 mm 的最后一个用户离开后,unmap 全部 VMA,再释放页表页25:

// mm/mmap.c:1273-1313
void exit_mmap(struct mm_struct *mm)
{
	...
	unmap_vmas(&tlb, &unmap);
	...
	free_pgtables(&tlb, &unmap);
	tlb_finish_mmu(&tlb);

物理页不是「进程一退就全扔」。文件页、共享库、还被别的 mm 指着的 COW 页,refcount 掉到 0 才回 buddy。这更像「丢掉这棵页表」,而不是格式化整台机器的内存。

内核没有 git log 那种地址空间时间旅行。fork 出来的 COW 页是分叉,不是同一份 mm 的快照链。

Git:reflog 可查,gc 清孤儿

commit 的 parents 就是历史链。git log 顺着它走。失去所有 ref / reflog 引用的对象,才是 gc 的对象。git gc 会跑 prune-packed,把已经打进 pack 的松散对象删掉26:

// builtin/gc.c:973-983
static int prune_packed(struct maintenance_run_opts *opts)
{
	struct child_process child = CHILD_PROCESS_INIT;

	child.git_cmd = 1;
	strvec_push(&child.args, "prune-packed");
	...
	return !!run_command(&child);
}

对象不可变,所以回收很朴素:还被 ref 指着的留下,没人指的删除。

Iceberg:time travel 与 expireSnapshots

snapshot 带 parentId() 和 timestampMillis(),规范里的 snapshot references 就是 branch / tag27。expireSnapshots 先提交一份去掉过期 snapshot 的新 metadata,再按引用关系删文件:

// core/src/main/java/org/apache/iceberg/RemoveSnapshots.java:360-379
  public void commit() {
    Tasks.foreach(ops)
        ...
            item -> {
              TableMetadata updated = internalApply();
              ops.commit(base, updated);
            });
    ...
    if (CleanupLevel.NONE != cleanupLevel && !base.snapshots().isEmpty()) {
      cleanExpiredSnapshots();
    }
  }

参见 RemoveSnapshots.java。和 Git 一样:先让指针和引用集合不再指向旧 snapshot,再物理删 Parquet / manifest。还被别的 branch、tag 或保留窗口钉住的文件不能删。


五、LVM:extent 上的同一套手法

逻辑卷也是「逻辑地址对物理块」。颗粒换成 extent,默认 4 MiB。LV 里的逻辑 extent 和 VG 里的物理 extent 一一对应28:

The logical extents within the LV correspond one-to-one with physical extents in the VG.

读一块 LV,是查这张映射,落到某块 PV 上的一段 PE。内核侧由 device-mapper 接着走。不是扫整块盘。

VG 的「当前状态」也是一枚指针。元数据是 ASCII,放在 PV 上的循环缓冲里。新配置先追加,再改指向这份文本的指针29:

A metadata area is a circular buffer. New metadata is appended to the old metadata and then the pointer to the start of it is updated.

如上所示,这和 Iceberg 写新 metadata.json 再换 metadata_location、Git 写完对象再 rename ref,是同一类提交。旧副本还能留在缓冲里。

快照要分两代。旧式 lvcreate -s 是 origin 一写就把旧块拷进 COW 区。thin pool 更像页表和 Git:快照先共享数据块,写才拆开。内核文档把共享写成这一代的卖点30:

it allows many virtual devices to be stored on the same data volume. This simplifies administration and allows the sharing of data between volumes, thus reducing disk usage.

dm-thin 把映射放在一棵 copy-on-write btree 里。打内部快照就是克隆根节点,之后没有「正本 / 副本」之分,只是两棵树碰巧指向同一批数据块:

// drivers/md/dm-thin.c:54-62
 * We use a standard copy-on-write btree to store the mappings for the
 * devices (note I'm talking about copy-on-write of the metadata here, not
 * the data).  When you take an internal snapshot you clone the root node
 * of the origin btree.  After this there is no concept of an origin or a
 * snapshot.  They are just two device trees that happen to point to the
 * same data blocks.

参见 dm-thin.c。写共享块时,数据落到新块,再插进「这一棵」树;另一棵的节点不动。注释说上次 commit 之后的 origin btree 原样留着,是函数式里的持久化数据结构。崩溃时两边都还指着旧块,看不见半成品。

可以看到,LVM 的叶子没有 Git blob 那么「写完就不改」。普通 LV 的 extent 可以原地写,像独占页走 wp_page_reuse()。真正靠新索引共享旧块的,是 thin 快照窗口。

flowchart LR
    MDA["MDA header"] --> Meta["VG metadata"]
    Meta --> LV["LV:LE → PE"]
    LV --> PE["PV 上的物理 extent"]

    style MDA fill:#87CEEB,stroke:#333,stroke-width:2px

对照一下:

维度 Linux 分页 Git Iceberg LVM
可变指针 CR3 / mm->pgd HEAD → refs/heads/* Catalog 的 metadata 路径,或 Hadoop 的 vN.metadata.json MDA header 指向当前 VG metadata
底层颗粒 物理页(页框) blob data file(常为 Parquet) 物理 extent(PE)
索引层 PGD … PTE commit → tree snapshot → manifest list → manifest LE → PE;thin 用 mapping btree
共享 不同页表可指向同一页 不同 commit 可指向同一 blob 不同 snapshot / manifest 可列出同一文件 thin 快照共享同一批数据块
提交 set_pte_at / write_cr3 lockfile + rename ref rename 或 CAS metadata_location 追加新 metadata,再改指针
回收 exit_mmap + 页 refcount git gc / prune expireSnapshots 后再删文件 删 LV / 拆共享后再还 PE

收成一句:底下是可共享的颗粒,上面叠一层或多层索引;换「当前状态」只换根上的指针,不搬叶子。

不同进程的页表可以指向同一张物理页(fork 之后、COW 之前);不同 commit 的 tree 可以指向同一个 blob;新 snapshot 的 manifest 会原样列出没改过的旧 data file;thin 快照的两棵 mapping btree 可以指向同一批数据块。叶子可以继续被旧根引用。索引叠几层,是为了把「找一块」从扫全表收成按路径走,并让共享发生在合适的粒度上。

颗粒怎么命名不一样:页框用 PFN,blob 用内容 hash,data file 用路径,PE 用 PV 上的偏移。颗粒的「不变」程度也不一样。blob 和已提交的 data file 写完就不改;物理页和普通 LV 的 extent 只在被共享、只读时当成这种叶子,独占后可以原地写。Iceberg 这边也不是「不同 manifest 必须指向不同 Parquet」——没改的文件就是被新索引接着指。

这套设计换来的能力可以对照看:

能力 Linux 分页 Git Iceberg LVM
状态切换便宜 write_cr3,不搬页 改 HEAD / checkout 换 Catalog 里的 metadata 路径 改 MDA 指针,不搬 PE
改一点不拷全部 fork 共享只读页 新 commit 复用旧 blob 新 snapshot 复用旧 data file thin 快照共享数据块
读不受写打扰 先 set_pte_at,读者走旧映射 先写对象再 update-ref 先写文件再 CAS / rename 先追加 metadata 再改指针;thin 写新块
还能回到旧根 fork 后的共享页是瞬时快照(没有地址空间日志) git log 沿 parent time travel,旧 snapshot 仍在 thin 快照;MDA 缓冲里的旧文本
查找能剪枝 按虚拟地址逐级走页表 沿路径走 tree manifest list / 列 bounds,再进 Parquet footer 按 LE 查到 PE
逻辑名和物理块脱钩 VA 对 PFN 文件名在 tree,内容在 blob 路径在 manifest,行在 data file LV 偏移对 PE
回收可以往后放 页 refcount git gc / prune expireSnapshots 后再删文件 删 LV / 拆共享后再还 PE

用间接层换来廉价快照、隔离写入、共享历史和可剪枝的查找,代价是多一层指针,以及必须另做垃圾回收。

LSM-Tree 也可以用同一副眼镜看:WAL / memtable 是新影子,SST 是不可变颗粒,compaction 是后台重写索引,manifest 是那枚指针。那是另一篇的事。

References

  1. Linux 内核 arch/x86/mm/tlb.c — switch_mm_irqs_off() / load_new_mm_cr3()。进程换 mm 时把 next->pgd 写入 CR3;注释强调 load_cr3() 的串行化语义。 ↩

  2. Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 3A: System Programming Guide, Part 1(Order 253668),Chapter 4 Paging。§4.5 四级/五级的开启条件;§4.2 从 CR3 起每次切 9 位的迭代走表。全套手册入口:Intel SDM。 ↩ ↩2 ↩3 ↩4

  3. Git 手册 git-update-ref — Update the object name stored in a ref safely。源码见 Documentation/git-update-ref.adoc。给定 <new-oid> <old-oid> 时先验证再写;files backend 用 lockfile + commit_lock_file() 完成单条 ref 的原子更新。 ↩

  4. Apache Iceberg Table Spec(源码 format/spec.md)— Overview / Optimistic Concurrency / File System Operations。表状态变更写新 metadata,并以原子交换替换旧指针;数据文件写后不可变;Hadoop 表用 rename 提交 metadata。 ↩ ↩2 ↩3 ↩4

  5. Apache Iceberg HadoopTableOperations.java — commit() / renameToFinal() / writeVersionHint()。注释写明 rename 是原子提交;version-hint.text 为 best-effort。 ↩

  6. Apache Iceberg BaseMetastoreTableOperations.java 中的 METADATA_LOCATION_PROP;HMSTablePropertyHelper.java 把新路径写入 HMS 表参数。 ↩

  7. Linux 内核文档 5-level paging(源码 Documentation/arch/x86/x86_64/5level-paging.rst)— a straight-forward extension of the current page table structure adding one more layer of translation。pgdir_shift 默认 39、五级改为 48,见 arch/x86/boot/compressed/pgtable_64.c。 ↩ ↩2

  8. 同上,§4.5.2 Use of CR3 with Ordinary 4-Level Paging and 5-Level Paging;Table 4-12(CR4.PCIDE = 0)bits 12 及以上为 4K 对齐的 PML4/PML5 物理地址。Linux 写入见 arch/x86/include/asm/special_insns.h 的 native_write_cr3()。 ↩

  9. 同上,§4.7 Page-Fault Exceptions;Figure 4-12 的 error code(P / W/R / U/S / RSVD)。Linux 对位的注释见 arch/x86/include/asm/trap_pf.h;入口 arch/x86/mm/fault.c 的 exc_page_fault() / do_user_addr_fault()。 ↩

  10. Linux 内核 include/linux/pgtable.h — pgd_offset() / pmd_off()。软件页表行走的折叠路径。 ↩

  11. Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 3C: System Programming Guide, Part 3(Order 326019-090US,February 2026),§27.4 Guest-State Area、§27.4.1 Guest Register State。VM entry 从这些字段装入处理器状态;guest-state 含 CR0、CR3、CR4。 ↩

  12. Volume 3A §4.1.2 Paging-Mode Enabling:VMX transitions allow transitions between paging modes that are not possible using MOV to CR or WRMSR. This is because VMX transitions can load CR0, CR4, and IA32_EFER in one operation. HLAT 用 VMCS 的 HLATP 定位第一张表,见 §4.5.3;细节在 Volume 3C 的 VMCS / VM-execution control。 ↩

  13. Prem Kumar Ponuthorai、Jon Loeliger,Version Control with Git 第 3 版(O’Reilly,2022),第 2 章 Foundational Concepts(「Blob Objects and Hashes」「Tree Object and Files」「Commit Objects」)。hello world blob 3b18e512dba79e4c8300dd08aeb37f8e728b8dad、tree 68aba62e560c0ebc3396e8ae9335232cd93a3f60,以及 A blob holds a file’s data but does not contain any metadata about the file or even its name。 ↩ ↩2 ↩3 ↩4 ↩5

  14. Tomer Shiran、Jason Hughes、Alex Merced,Apache Iceberg: The Definitive Guide(O’Reilly,2024),第 2 章 The Architecture of Apache Iceberg。data layer 的 data file / delete file;「file format most commonly used is Apache Parquet」;湖上文件按不可变处理。 ↩ ↩2

  15. Git object.h — enum object_type;commit.h — struct commit。 ↩

  16. Apache Iceberg spec Manifests — A manifest is an immutable Avro file that lists data files or delete files;这些 manifest 由每个 snapshot 的 manifest list 跟踪。 ↩

  17. Apache Iceberg spec Scan Planning(源码 format/spec.md#scan-planning);实现见 ManifestEvaluator.java、InclusiveMetricsEvaluator.java。 ↩

  18. Apache Parquet File format;仓库说明 README.md#file-format;Thrift FileMetaData。读者先读 footer,再按 row group / column chunk 定位。 ↩

  19. Apache Iceberg 测试夹具 core/src/test/resources/TableMetadataV2Valid.json。current-snapshot-id、parent-snapshot-id、manifest-list 的官方样例。 ↩

  20. Apache Iceberg TestReachableFileUtil.java — DataFiles.builder(SPEC).withPath("/path/to/data-a.parquet");路径 API 见 DataFiles.java。 ↩

  21. Apache Iceberg spec Data File Fields — file_path(字段 100)、file_format(101)、record_count(103)。manifest 为 Avro,文中 JSON 只用来对照字段。 ↩

  22. Linux 内核 mm/memory.c — do_wp_page()。私有映射写过错:能复用则 wp_page_reuse(),否则 wp_page_copy()。 ↩

  23. Linux 内核 mm/memory.c — wp_page_copy() 里 ptep_clear_flush() 之后才 set_pte_at(),并说明必须先切换 PTE 再减旧页 mapcount。 ↩

  24. Apache Iceberg SnapshotProducer.java — commit()。apply() 出新 snapshot,再 ops.commit(base, updated);只对 CommitFailedException 按表属性重试。 ↩

  25. Linux 内核 mm/mmap.c — exit_mmap()。最后一个 mm 用户离开后 unmap VMA 并 free_pgtables()。 ↩

  26. Git builtin/gc.c — prune_packed()。 ↩

  27. Apache Iceberg Snapshot.java;spec Snapshot References;回收见 RemoveSnapshots.java。 ↩

  28. Red Hat Enterprise Linux 9,Managing LVM volume groups — Extents are the smallest units of space that you can allocate in LVM;默认 4 MiB;The logical extents within the LV correspond one-to-one with physical extents in the VG。 ↩

  29. Red Hat Enterprise Linux 7,Appendix E. LVM Volume Group Metadata — A metadata area is a circular buffer. New metadata is appended to the old metadata and then the pointer to the start of it is updated. 元数据为 ASCII,默认在每个 PV 的 metadata area 留一份拷贝。 ↩

  30. Linux 内核文档 Thin provisioning(源码 Documentation/admin-guide/device-mapper/thin-provisioning.rst)— 多份虚拟设备共享同一 data volume。实现见 drivers/md/dm-thin.c:内部快照克隆 mapping btree 的根节点;写共享块时数据落到新块。 ↩