CR3、HEAD 与 Catalog:分页、Git、Iceberg 共用的「可变指针 + 不可变树」
最近在对照 Linux 分页、Git 对象模型和 Iceberg 表格式,发现它们在解同一道题:底层存储块一旦写下就不改,逻辑状态却要持续变。答案都是「一枚可变指针 + 一棵不可变树」。在这里把三套源码并排看一下。
一、当前状态只是一枚指针
共通做法是:数据本身不可变,真正会改的只有「现在指向谁」。切换可见状态,就是改这枚指针。
Linux:CR3 指向页表根
x86-64 上,CR3 存的是当前地址空间顶层页表(PGD)的物理地址,外加 PCID 等控制位。切到另一个进程的 mm_struct 时,内核把新的 next->pgd 装进 CR3:
// arch/x86/mm/tlb.c:565-582
static void load_new_mm_cr3(pgd_t *pgdir, u16 new_asid, unsigned long lam,
bool need_flush)
{
unsigned long new_mm_cr3;
if (need_flush) {
invalidate_user_asid(new_asid);
new_mm_cr3 = build_cr3(pgdir, new_asid, lam);
} else {
new_mm_cr3 = build_cr3_noflush(pgdir, new_asid, lam);
}
/*
* Caution: many callers of this function expect
* that load_cr3() is serializing and orders TLB
* fills with respect to the mm_cpumask writes.
*/
write_cr3(new_mm_cr3);
}
参见 arch/x86/mm/tlb.c。switch_mm_irqs_off() 在真正换 mm 时走到这里1:
// arch/x86/mm/tlb.c:947-956
if (ns.need_flush) {
VM_WARN_ON_ONCE(is_global_asid(ns.asid));
this_cpu_write(cpu_tlbstate.ctxs[ns.asid].ctx_id, next->context.ctx_id);
this_cpu_write(cpu_tlbstate.ctxs[ns.asid].tlb_gen, next_tlb_gen);
load_new_mm_cr3(next->pgd, ns.asid, new_lam, true);
可以看到,进程切换并不搬页表,也不拷物理页。它改的是 CPU 手里那枚「当前页表根」指针。write_cr3 还是一条串行化指令,后面的 TLB 填充会按新根走。CR3 里放什么、CPU 怎么用它走路,写在 Intel SDM Volume 3A 第 4 章2;内核按这个格式填表,再 mov 进寄存器。
Git:HEAD 先指分支,分支再指 commit
.git/HEAD 通常不是 commit hash,而是一条符号引用,例如 ref: refs/heads/main。真正存 hash 的是分支文件。detach 时 HEAD 才直接写 object id。git commit 先写出新的 commit 对象,再把当前分支指针推过去:
// builtin/commit.c:1938-1946
if (commit_tree_extended(sb.buf, sb.len, &the_repository->index->cache_tree->oid,
parents, &oid, author_ident.buf, NULL,
sign_commit, extra)) {
rollback_index_files();
die(_("failed to write commit object"));
}
if (update_head_with_reflog(current_head, &oid, reflog_msg, &sb,
&err)) {
参见 builtin/commit.c。git-update-ref 把这件事说得很干净:给定新旧 oid,验证旧值后再写入新值3。files backend 先把 hex oid 写进 lockfile,再 rename 成正式 ref:
// refs/files-backend.c:2059-2077
static enum ref_transaction_error write_ref_to_lockfile(struct files_ref_store *refs,
struct ref_lock *lock,
const struct object_id *oid,
struct strbuf *err)
{
static char term = '\n';
int fd;
fd = get_lock_file_fd(&lock->lk);
if (write_in_full(fd, oid_to_hex(oid), refs->base.repo->hash_algo->hexsz) < 0 ||
write_in_full(fd, &term, 1) < 0 ||
fsync_component(FSYNC_COMPONENT_REFERENCE, get_lock_file_fd(&lock->lk)) < 0 ||
close_ref_gently(lock) < 0) {
// refs/files-backend.c:1864-1892
static int commit_ref(struct ref_lock *lock)
{
char *path = get_locked_file_path(&lock->lk);
...
if (commit_lock_file(&lock->lk))
return -1;
return 0;
}
如上所示,读者看见的「当前分支」,始终是那枚 ref 指针;对象库里的 commit/tree/blob 写完就不会改。
Iceberg:Catalog 存 metadata 路径
Iceberg 规范把表状态放在 metadata 文件里,每次变更都写一份新文件,再用原子交换替换旧指针4:
All changes to table state create a new metadata file and replace the old metadata with an atomic swap.
指针落在哪,取决于 Catalog 实现。
Hadoop 路径表没有外部 metastore。原子提交是把临时 metadata rename 成下一版本号文件(vN.metadata.json)。源码把这句话写在注释里:
// core/src/main/java/org/apache/iceberg/hadoop/HadoopTableOperations.java:157-167
int nextVersion = (current.first() != null ? current.first() : 0) + 1;
Path finalMetadataFile = metadataFilePath(nextVersion, codec);
FileSystem fs = getFileSystem(tempMetadataFile, conf);
// this rename operation is the atomic commit operation
renameToFinal(fs, tempMetadataFile, finalMetadataFile, nextVersion);
LOG.info("Committed a new metadata file {}", finalMetadataFile);
// update the best-effort version pointer
writeVersionHint(nextVersion);
参见 HadoopTableOperations.java。version-hint.text 只是 best-effort 加速查找,丢了可以扫目录恢复;真正互斥的是「vN 这份文件是否已经存在」。renameToFinal() 发现目标已在,就当并发提交失败5。
Hive / JDBC 这类 metastore Catalog,指针是表属性里的 metadata_location:
// core/src/main/java/org/apache/iceberg/BaseMetastoreTableOperations.java:46-50
public static final String TABLE_TYPE_PROP = "table_type";
public static final String ICEBERG_TABLE_TYPE_VALUE = "iceberg";
public static final String METADATA_LOCATION_PROP = "metadata_location";
public static final String METADATA_HASH_PROP = "metadata_hash";
public static final String PREVIOUS_METADATA_LOCATION_PROP = "previous_metadata_location";
Hive 提交前先核对「我看到的 base 路径」是不是 HMS 里当前那条,对不上就拒掉:
// hive-metastore/.../HiveTableOperations.java:304-310
String metadataLocation = tbl.getParameters().get(METADATA_LOCATION_PROP);
String baseMetadataLocation = base != null ? base.metadataFileLocation() : null;
if (!Objects.equals(baseMetadataLocation, metadataLocation)) {
throw new CommitFailedException(
"Cannot commit: Base metadata location '%s' is not same as the current table metadata location '%s' for %s.%s",
baseMetadataLocation, metadataLocation, database, tableName);
参见 HiveTableOperations.java。通过之后再把 metadata_location 写成新文件路径6。读者刷新 Catalog,拿到的就是新的 metadata.json。
三套系统并排看:
flowchart LR
subgraph Linux["Linux x86"]
CR3["CR3"] --> PGD["mm.pgd / 页表根"]
end
subgraph GitBox["Git"]
HEAD["HEAD"] --> Branch["refs/heads/*"]
Branch --> Commit["commit object"]
end
subgraph IcebergBox["Iceberg"]
Cat["Catalog / vN.metadata.json"] --> Meta["metadata.json"]
Meta --> Snap["current-snapshot-id"]
end
style CR3 fill:#87CEEB,stroke:#333,stroke-width:2px
style HEAD fill:#87CEEB,stroke:#333,stroke-width:2px
style Cat fill:#87CEEB,stroke:#333,stroke-width:2px
二、多层漏斗:按图索骥,不扫整片海
指针只解决「当前是哪一棵树」。树本身还得能快速缩小范围,否则每次访问都要遍历全部底层小块。这里的复杂度差就是 O(log N) / 分区裁剪 vs O(N) 全扫。
页表:PGD → P4D → PUD → PMD → PTE
树的形状是硬件定的。Intel SDM Volume 3A 第 4 章写分页2;Linux 按那套格式在内存里建表。CPU 查表时不跑内核。
四级还是五级,看四枚控制位2:
A logical processor uses 4-level paging if CR0.PG = 1, CR4.PAE = 1, IA32_EFER.LME = 1, and CR4.LA57 = 0. 4-level paging translates 48-bit linear addresses to 52-bit physical addresses.
A logical processor uses 5-level paging if CR0.PG = 1, CR4.PAE = 1, IA32_EFER.LME = 1, and CR4.LA57 = 1. 5-level paging translates 57-bit linear addresses to 52-bit physical addresses.
走表是一次迭代。第一张表的物理地址在 CR3 里;线性地址每次切若干位选一项,项要么指向下一张表,要么指向页框2:
The first paging structure used for any translation is located at the physical address in CR3.
With 4-level paging, each paging structure comprises 512 entries and translation uses 9 bits at a time from a 48-bit linear address. Bits 47:39 identify the first paging-structure entry, bits 38:30 identify a second, bits 29:21 a third, and bits 20:12 identify a fourth.
5-level paging is similar to 4-level paging except that 5-level paging translates 57-bit linear addresses. Bits 56:48 identify the first paging-structure entry, while the remaining bits are used as with 4-level paging.
Linux 把头文件写成同一组位移。每一级 512 项;P4D_SHIFT 39、PUD_SHIFT 30、PMD_SHIFT 21,就是上面那几组 9 位:
// arch/x86/include/asm/pgtable_64_types.h:50-80
#define PGDIR_SHIFT pgdir_shift
#define PTRS_PER_PGD 512
#define P4D_SHIFT 39
#define MAX_PTRS_PER_P4D 512
#define PTRS_PER_P4D ptrs_per_p4d
#define PUD_SHIFT 30
#define PTRS_PER_PUD 512
#define PMD_SHIFT 21
#define PTRS_PER_PMD 512
#define PTRS_PER_PTE 512
参见 pgtable_64_types.h。选哪一项,就是右移再掩码:
// include/linux/pgtable.h:69-72
#define pgd_index(a) (((a) >> PGDIR_SHIFT) & (PTRS_PER_PGD - 1))
pgdir_shift 默认 39,顶层对应手册的 PML4(线性地址 bits 47:39)。打开五级后改成 48,顶层变成 PML5(bits 56:48)7:
// arch/x86/boot/compressed/pgtable_64.c:16-17,121-128
unsigned int __section(".data") pgdir_shift = 39;
...
if (!cmdline_find_option_bool("no5lvl") &&
native_cpuid_eax(0) >= 7 && (native_cpuid_ecx(7) & BIT(16))) {
l5_required = true;
__pgtable_l5_enabled = 1;
pgdir_shift = 48;
ptrs_per_p4d = 512;
}
如上所示,CPUID.(EAX=07H,ECX=0):ECX.LA57 那一位,手册 §4.1.4 和压缩内核里查的是同一位。内核文档把五级写成「在现有页表上再加一层」7。
CR3 的用法手册单独写了一节8:
Ordinary 4-level paging and 5-level paging each translate linear addresses using a hierarchy of in-memory paging structures located using the contents of CR3, which is used to locate the first paging structure. For 4-level paging, this is the PML4 table, and for 5-level paging it is the PML5 table.
Table 4-12:CR4.PCIDE = 0 时,bits 12 及以上是 4K 对齐的 PML4/PML5 物理地址。Linux 切 mm 时 load_new_mm_cr3() 算出新值,最后就是一条 mov 进 CR3:
// arch/x86/include/asm/special_insns.h:41-44
static __always_inline void native_write_cr3(unsigned long val)
{
asm volatile("mov %0,%%cr3": : "r" (val) : "memory");
}
参见 special_insns.h。
项不存在、或保留位置了 1,硬件不翻译,直接 #PF(异常 14)9:
Accesses using linear addresses may cause page-fault exceptions (#PF; exception 14). An access to a linear address may cause a page-fault exception for either of two reasons: (1) there is no translation for the linear address; or (2) there is a translation for the linear address, but its access rights do not permit the access.
there is no translation for a linear address if the translation process for that address would use a paging-structure entry in which the P flag (bit 0) is 0 or one that sets a reserved bit.
出错线性地址在 CR2。Linux 从这条向量进来,读 CR2,再决定建表、杀进程还是修 PTE:
// arch/x86/mm/fault.c:1483-1488
DEFINE_IDTENTRY_RAW_ERRORCODE(exc_page_fault)
{
irqentry_state_t state;
unsigned long address;
address = cpu_feature_enabled(X86_FEATURE_FRED) ? fred_event_data(regs) : read_cr2();
// arch/x86/include/asm/trap_pf.h:7-25
* bit 0 == 0: no page found 1: protection fault
* bit 1 == 0: read access 1: write access
* bit 2 == 0: kernel-mode access 1: user-mode access
* bit 3 == 1: use of reserved bit detected
...
X86_PF_PROT = BIT(0),
X86_PF_WRITE = BIT(1),
X86_PF_USER = BIT(2),
X86_PF_RSVD = BIT(3),
可以看到,这些位就是手册 Figure 4-12 的 P / W/R / U/S / RSVD。用户地址落到 do_user_addr_fault(),再 handle_mm_fault()。内核这才按树往下走,缺哪一级就分配哪一级:
// mm/memory.c:6466-6549
pgd = pgd_offset(mm, address);
p4d = p4d_alloc(mm, pgd, address);
if (!p4d)
return VM_FAULT_OOM;
vmf.pud = pud_alloc(mm, p4d, address);
...
vmf.pmd = pmd_alloc(mm, vmf.pud, address);
...
fallback:
return handle_pte_fault(&vmf);
参见 mm/memory.c。已知地址已经落到 PMD 时,还有一条折叠助手,把四级偏移写成一行10:
// include/linux/pgtable.h:165-168
static inline pmd_t *pmd_off(struct mm_struct *mm, unsigned long va)
{
return pmd_offset(pud_offset(p4d_offset(pgd_offset(mm, va), va), va), va);
}
Volume 3C 不重写这套走表。它把同一组寄存器收进 VMCS 的 guest-state,进 guest 时装上去11:
VM entries load processor state from these fields and VM exits store processor state into these fields.
Control registers CR0, CR3, and CR4 (64 bits each; 32 bits on processors that do not support Intel 64 architecture).
guest 里的线性地址翻译,还是 3A 第 4 章。3A 自己点到了虚拟化:VMX 可以一次装入 CR0、CR4、IA32_EFER,切到普通 MOV / WRMSR 做不到的分页模式;HLAT 用 VMCS 里的 HLATP 代替 CR3 当第一张表12。3C 叠的是 VMM 对 CR3 的保存/恢复,不是另一套用户页表格式。
Version Control with Git 里的 blob / tree / commit、Apache Iceberg: The Definitive Guide 里的 snapshot / manifest,是软件自己定的树1314。页表这棵树,格式先写在 SDM 里;Linux 的 pgd_t / pte_t 是按那份格式填的格子。没有 Linux,别的 OS 也得按同一套 CR3 跟 CPU 打交道。MMU 不会扫进程的全部物理页,只按虚拟地址切出每一级索引。
Git:commit → tree → blob
对象类型就三种常用的(再加上 tag)15:
// object.h:99-104
enum object_type {
OBJ_BAD = -1,
OBJ_NONE = 0,
OBJ_COMMIT = 1,
OBJ_TREE = 2,
OBJ_BLOB = 3,
OBJ_TAG = 4,
commit 只记住一棵根 tree,以及 parent 链:
// commit.h:27-39
struct commit {
struct object object;
timestamp_t date;
struct commit_list *parents;
struct tree *maybe_tree;
unsigned int index;
};
tree 是目录项列表,项指向下一层 tree 或 blob。git commit 用 index 上的 cache_tree 当根 tree oid,不会为没改过的子树重写对象。查一个文件是「沿路径走目录项」,不是枚举整个对象库。
Iceberg:metadata → snapshot → manifest list → manifest → Parquet
规范里的快照结构是4:
- metadata.json 记下 schema、partition spec,以及
current-snapshot-id - 每个 snapshot 有一份 manifest list
- manifest list 里是若干 manifest,带分区统计和文件计数
- manifest 里才是 data file / delete file 路径和列度量
- 数据文件本身通常是 Parquet(也可以是 Avro / ORC)
Parquet 文件列表不在 metadata.json 里。规范把 manifest 写成一份不可变 Avro:列出 data file 或 delete file,以及分区、度量和跟踪信息;一个 snapshot 的这些 manifest 再由 manifest list 索引一层16:
Catalog
└─ metadata.json 当前 snapshot id,不列数据文件
└─ snapshot
└─ manifest list 列出有哪些 manifest(分区摘要、文件计数)
└─ manifest ← data file / delete file 清单在这里
└─ /path/to/data-a.parquet
metadata.json 只回答「当前是哪次 snapshot」,以及这次 snapshot 的 manifest list 路径。manifest list 回答「哪些 manifest 值得打开」。真正的 file_path 在 manifest 的 data_file 里。一份 snapshot 通常有多份 manifest,每份只覆盖一部分文件。
Java 里当前快照就是按 id 取:
// core/src/main/java/org/apache/iceberg/TableMetadata.java:536-538
public Snapshot currentSnapshot() {
return snapshotsById.get(currentSnapshotId);
}
扫描规划明确写了可以跳过整份 manifest17:
Manifests that contain no matching files, determined using either file counts or partition summaries, may be skipped.
实现上,ManifestEvaluator 用 manifest 的分区摘要判断这份文件里有没有可能命中的分区;过了这一关,再用 InclusiveMetricsEvaluator 看单个 data file 的列上下界。eval 返回 false,这份文件就可以不读。
叶子上的 Parquet 自己又是一层漏斗。文件尾部是 FileMetaData,里面是 row group 列表;读者先读 footer,再只打开关心的 column chunk / page18:
4-byte magic number "PAR1"
<Column chunks / row groups>
File Metadata
4-byte length of file metadata
4-byte magic number "PAR1"
struct FileMetaData {
1: required i32 version
2: required list<SchemaElement> schema
3: required i64 num_rows
4: required list<RowGroup> row_groups
...
}
参见 parquet.thrift。Iceberg 决定读哪些文件,Parquet footer 再决定读文件里的哪些列、哪些 row group。
三棵树并排:
flowchart TB
subgraph PT["页表"]
VA["虚拟地址"] --> PGD2["PGD"]
PGD2 --> P4D["P4D"]
P4D --> PUD["PUD"]
PUD --> PMD["PMD"]
PMD --> PTE["PTE / 物理页"]
end
subgraph GT["Git"]
BR["branch / HEAD"] --> CM["commit"]
CM --> TR["tree"]
TR --> TR2["tree"]
TR --> BL["blob"]
TR2 --> BL2["blob"]
end
subgraph IB["Iceberg + Parquet"]
CAT["Catalog"] --> MD["metadata.json"]
MD --> SN["snapshot"]
SN --> ML["manifest list"]
ML --> MF["manifest"]
MF --> PQ["Parquet file"]
PQ --> FT["footer / row group / page"]
end
叶子这一层,Git 的 blob 对上 Iceberg 的 data file。data file 常常是 Parquet,但表格式并不绑定它。Loeliger / Ponuthorai 那本书把 blob 写成不透明字节,连文件名都不在对象里;路径在 tree 上13。Iceberg 那本书把叶子放在 data layer:manifest 跟踪文件,行数据在 data file 里,Parquet 只是最常见的一种封装14。
下面用两边都能核对的样例走一遍。
Git 书第 2 章这份 12 字节加换行,hash 是固定的13:
$ echo "hello world" > hello.txt
$ git add hello.txt
$ echo "hello world" | git hash-object --stdin
3b18e512dba79e4c8300dd08aeb37f8e728b8dad
$ git cat-file -p 3b18e512dba79e4c8300dd08aeb37f8e728b8dad
hello world
对象落在 .git/objects/3b/18e512dba79e4c8300dd08aeb37f8e728b8dad。git add 只把内容和 pathname 记进 index,还没有 tree:
$ git ls-files -s
100644 3b18e512dba79e4c8300dd08aeb37f8e728b8dad 0 hello.txt
git write-tree 才把「名字 → blob」冻成 tree13:
$ git write-tree
68aba62e560c0ebc3396e8ae9335232cd93a3f60
$ git cat-file -p 68aba6
100644 blob 3b18e512dba79e4c8300dd08aeb37f8e728b8dad hello.txt
如上所示,hello.txt 这个名字在 tree 里,不在 blob 里。再 commit 一次,commit 对象只多记 tree、作者和时间;书里的例子是13:
tree 492413269336d21fac079d4a4672e55d5d2147ac
author Jon Loeliger <jdl@example.com> 1656932750 +0200
committer Jon Loeliger <jdl@example.com> 1656932750 +0200
Commit a file that says hello
commit hash 会因作者和时间而变,tree 可以原样复用。HEAD 通常还是符号引用:
ref: refs/heads/main
refs/heads/main 里才是那串 commit oid。
Iceberg 官方测试夹具 TableMetadataV2Valid.json 把同一角色写成 JSON。Catalog 指向这份 metadata;当前可见状态是 current-snapshot-id19:
{
"format-version": 2,
"table-uuid": "9c12d441-03fe-4693-9a96-a0705ddf69c1",
"location": "s3://bucket/test/location",
"current-snapshot-id": 3055729675574597004,
"snapshots": [
{
"snapshot-id": 3051729675574597004,
"timestamp-ms": 1515100955770,
"sequence-number": 0,
"summary": { "operation": "append" },
"manifest-list": "s3://a/b/1.avro"
},
{
"snapshot-id": 3055729675574597004,
"parent-snapshot-id": 3051729675574597004,
"timestamp-ms": 1555100955770,
"sequence-number": 1,
"summary": { "operation": "append" },
"manifest-list": "s3://a/b/2.avro",
"schema-id": 1
}
]
}
parent-snapshot-id 对上 Git commit 的 parent。manifest-list 对上 commit 里的 tree:下一层索引的位置,不是行数据本身。
再往下一层,测试代码用 DataFiles.builder 造叶子。路径、大小、行数写在 builder 上,不按内容算 hash20:
// core/src/test/java/org/apache/iceberg/util/TestReachableFileUtil.java:58-63
private static final DataFile FILE_A =
DataFiles.builder(SPEC)
.withPath("/path/to/data-a.parquet")
.withFileSizeInBytes(10)
.withRecordCount(1)
.build();
规范把这条记录展开成 manifest 里的 data_file struct。字段 id 是固定的:100 file_path、101 file_format、103 record_count、125 lower_bounds21。写成 JSON 看结构,就是:
{
"status": 1,
"snapshot_id": 3055729675574597004,
"data_file": {
"content": 0,
"file_path": "/path/to/data-a.parquet",
"file_format": "PARQUET",
"record_count": 1,
"file_size_in_bytes": 10
}
}
status = 1 是 ADDED。真正的 hello world 行在 Parquet 文件里,不在这条元数据里。读者要读内容,得按 file_path 打开文件,再读 footer 里的 row group。
同一份「hello world」并排看:
| 角色 | Git 样例 | Iceberg 样例 |
|---|---|---|
| 可变指针 | HEAD → ref: refs/heads/main |
Catalog 的 metadata_location → 上面这份 metadata.json |
| 当前根 | refs/heads/main 里的 commit oid |
"current-snapshot-id": 3055729675574597004 |
| 目录 / 文件列表 | tree 68aba62e…:100644 blob 3b18e5… hello.txt |
snapshot 的 manifest-list: s3://a/b/2.avro,再进 manifest |
| 叶子怎么命名 | 内容 SHA:3b18e512dba79e4c8300dd08aeb37f8e728b8dad |
路径:/path/to/data-a.parquet |
| 叶子里有没有名字 | blob 只有 hello world\n |
data file 的名字在 file_path;Parquet 里是列和行 |
| 拷一份同内容 | 第二个文件名仍指向同一个 blob | 另一条路径就是另一个 data file,即使字节相同 |
可以看到,说「blob 对应 Parquet」只在「不可变载荷叶子」这层成立。更贴的说法是 blob ↔ data file,tree ↔ manifest。Git 用内容做主键,所以两个路径可以共享一个 blob;Iceberg 用路径做主键,剪枝靠 manifest 上的分区和列上下界,不靠「这份 Parquet 的 SHA 跟另一份一样」。
三、先写影子,再原子切换
修改不能直接打在别人正在读的那份数据上。三套系统都是:在旁边做好新版本,最后一步才让指针看见它。
缺页分配物理页,和写时复制(COW)不是同一条路径。前者是「这页还没有」;后者是「这页有,但是只读共享,写就要拷一份」。内核把它们分成两条 fault 路径。
Linux:do_wp_page 拷页,再换 PTE
私有映射上的写过错到 do_wp_page()。能复用就复用;必须拷的时候走 wp_page_copy()22:
// mm/memory.c:4291-4320
if (folio && folio_test_anon(folio) &&
(PageAnonExclusive(vmf->page) || wp_can_reuse_anon_folio(folio, vma))) {
...
wp_page_reuse(vmf, folio);
return 0;
}
...
return wp_page_copy(vmf);
拷完之后,先清旧 PTE 并冲 TLB,再挂上新页。注释写得很清楚:必须先切换页表项,才能把旧页的 mapcount 减掉,否则别的进程可能在窗口里写进旧页23:
// mm/memory.c:3918-3929
ptep_clear_flush(vma, vmf->address, vmf->pte);
folio_add_new_anon_rmap(new_folio, vma, vmf->address, RMAP_EXCLUSIVE);
folio_add_lru_vma(new_folio, vma);
BUG_ON(unshare && pte_write(entry));
set_pte_at(mm, vmf->address, vmf->pte, entry);
对这个进程来说,写操作成功了;共享这份旧页的其他进程,页表还指着原来的只读页。可见性切换发生在这一条 PTE,不是整棵页表重写。
fork 之后父子共享只读页、一方先写再触发上面这条路径,就是教科书里的进程级 COW。进程切换本身仍然只写 CR3。
Git:对象先落盘,ref 后移动
commit_tree_extended() 先拼 commit 缓冲区,校验 tree 类型,再写入对象库:
// commit.c:1729-1760
int commit_tree_extended(const char *msg, size_t msg_len,
const struct object_id *tree,
const struct commit_list *parents, struct object_id *ret,
...)
{
...
odb_assert_oid_type(the_repository->objects, tree, OBJ_TREE);
...
write_commit_tree(&buffer, msg, msg_len, tree, parent_buf, nparents, author, committer, extra);
参见 commit.c。这一步失败,HEAD 不动。写成功之后,update_head_with_reflog() 才去锁 ref、写 lockfile、commit_lock_file() rename。并发更新用「期望的旧 oid」做 CAS,对不上就失败——和 Iceberg 核对 metadata_location 是同一类约束。
没改过的 blob / 子 tree 继续被新 tree 引用,这就是 Git 的 COW:只为变化路径分配新对象。
Iceberg:数据文件不可变,提交只换 metadata 指针
规范要求:文件写下去就不改;表不需要随机写。Hadoop 表才依赖 rename 实现 metadata 提交4。一次 append 大致是:
- 写出新的 Parquet(以及需要的 delete file)
- 写出新的 manifest;旧 snapshot 里还能用的 manifest 直接复用
- 写出新的 manifest list 和 metadata.json
- Catalog 原子替换指针
SnapshotProducer.commit() 先 apply() 得到新 snapshot,再 taskOps.commit(base, updated);撞上 CommitFailedException 就按 commit.num-retries 重试24:
// core/src/main/java/org/apache/iceberg/SnapshotProducer.java:480-522
public void commit() {
AtomicLong newSnapshotId = new AtomicLong(-1L);
...
taskOps -> {
Snapshot newSnapshot = apply();
newSnapshotId.set(newSnapshot.snapshotId());
TableMetadata.Builder update = TableMetadata.buildFrom(base);
...
TableMetadata updated = update.build();
if (updated.changes().isEmpty()) {
return;
}
taskOps.commit(base, updated.withUUID());
});
重试时序列号会重分,但新 manifest 可以复用——规范把这件事设计进了「从 manifest list 继承 sequence number」4。读者在指针切换前一直看着旧 snapshot,不会看见半成品文件。
sequenceDiagram
participant W as Writer
participant Shadow as 影子文件
participant Ptr as 指针CR3或HEAD或Catalog
participant R as 并发读者
R->>Ptr: 读当前指针
Ptr-->>R: 旧根
W->>Shadow: 写新页或新对象或新metadata
Note over W,Shadow: 读者仍走旧根
W->>Ptr: 原子切换
W-->>R: 下次刷新才看到新根
四、历史还在,回收另做
指针往前走以后,旧树不必立刻消失。历史查询靠「还有没有人引用」;物理回收是另一次显式动作。
Linux:进程退出拆掉页表
exit_mmap() 在 mm 的最后一个用户离开后,unmap 全部 VMA,再释放页表页25:
// mm/mmap.c:1273-1313
void exit_mmap(struct mm_struct *mm)
{
...
unmap_vmas(&tlb, &unmap);
...
free_pgtables(&tlb, &unmap);
tlb_finish_mmu(&tlb);
物理页不是「进程一退就全扔」。文件页、共享库、还被别的 mm 指着的 COW 页,refcount 掉到 0 才回 buddy。这更像「丢掉这棵页表」,而不是格式化整台机器的内存。
内核没有 git log 那种地址空间时间旅行。fork 出来的 COW 页是分叉,不是同一份 mm 的快照链。
Git:reflog 可查,gc 清孤儿
commit 的 parents 就是历史链。git log 顺着它走。失去所有 ref / reflog 引用的对象,才是 gc 的对象。git gc 会跑 prune-packed,把已经打进 pack 的松散对象删掉26:
// builtin/gc.c:973-983
static int prune_packed(struct maintenance_run_opts *opts)
{
struct child_process child = CHILD_PROCESS_INIT;
child.git_cmd = 1;
strvec_push(&child.args, "prune-packed");
...
return !!run_command(&child);
}
对象不可变,所以回收很朴素:还被 ref 指着的留下,没人指的删除。
Iceberg:time travel 与 expireSnapshots
snapshot 带 parentId() 和 timestampMillis(),规范里的 snapshot references 就是 branch / tag27。expireSnapshots 先提交一份去掉过期 snapshot 的新 metadata,再按引用关系删文件:
// core/src/main/java/org/apache/iceberg/RemoveSnapshots.java:360-379
public void commit() {
Tasks.foreach(ops)
...
item -> {
TableMetadata updated = internalApply();
ops.commit(base, updated);
});
...
if (CleanupLevel.NONE != cleanupLevel && !base.snapshots().isEmpty()) {
cleanExpiredSnapshots();
}
}
参见 RemoveSnapshots.java。和 Git 一样:先让指针和引用集合不再指向旧 snapshot,再物理删 Parquet / manifest。还被别的 branch、tag 或保留窗口钉住的文件不能删。
五、LVM:extent 上的同一套手法
逻辑卷也是「逻辑地址对物理块」。颗粒换成 extent,默认 4 MiB。LV 里的逻辑 extent 和 VG 里的物理 extent 一一对应28:
The logical extents within the LV correspond one-to-one with physical extents in the VG.
读一块 LV,是查这张映射,落到某块 PV 上的一段 PE。内核侧由 device-mapper 接着走。不是扫整块盘。
VG 的「当前状态」也是一枚指针。元数据是 ASCII,放在 PV 上的循环缓冲里。新配置先追加,再改指向这份文本的指针29:
A metadata area is a circular buffer. New metadata is appended to the old metadata and then the pointer to the start of it is updated.
如上所示,这和 Iceberg 写新 metadata.json 再换 metadata_location、Git 写完对象再 rename ref,是同一类提交。旧副本还能留在缓冲里。
快照要分两代。旧式 lvcreate -s 是 origin 一写就把旧块拷进 COW 区。thin pool 更像页表和 Git:快照先共享数据块,写才拆开。内核文档把共享写成这一代的卖点30:
it allows many virtual devices to be stored on the same data volume. This simplifies administration and allows the sharing of data between volumes, thus reducing disk usage.
dm-thin 把映射放在一棵 copy-on-write btree 里。打内部快照就是克隆根节点,之后没有「正本 / 副本」之分,只是两棵树碰巧指向同一批数据块:
// drivers/md/dm-thin.c:54-62
* We use a standard copy-on-write btree to store the mappings for the
* devices (note I'm talking about copy-on-write of the metadata here, not
* the data). When you take an internal snapshot you clone the root node
* of the origin btree. After this there is no concept of an origin or a
* snapshot. They are just two device trees that happen to point to the
* same data blocks.
参见 dm-thin.c。写共享块时,数据落到新块,再插进「这一棵」树;另一棵的节点不动。注释说上次 commit 之后的 origin btree 原样留着,是函数式里的持久化数据结构。崩溃时两边都还指着旧块,看不见半成品。
可以看到,LVM 的叶子没有 Git blob 那么「写完就不改」。普通 LV 的 extent 可以原地写,像独占页走 wp_page_reuse()。真正靠新索引共享旧块的,是 thin 快照窗口。
flowchart LR
MDA["MDA header"] --> Meta["VG metadata"]
Meta --> LV["LV:LE → PE"]
LV --> PE["PV 上的物理 extent"]
style MDA fill:#87CEEB,stroke:#333,stroke-width:2px
对照一下:
| 维度 | Linux 分页 | Git | Iceberg | LVM |
|---|---|---|---|---|
| 可变指针 | CR3 / mm->pgd |
HEAD → refs/heads/* |
Catalog 的 metadata 路径,或 Hadoop 的 vN.metadata.json |
MDA header 指向当前 VG metadata |
| 底层颗粒 | 物理页(页框) | blob | data file(常为 Parquet) | 物理 extent(PE) |
| 索引层 | PGD … PTE | commit → tree | snapshot → manifest list → manifest | LE → PE;thin 用 mapping btree |
| 共享 | 不同页表可指向同一页 | 不同 commit 可指向同一 blob | 不同 snapshot / manifest 可列出同一文件 | thin 快照共享同一批数据块 |
| 提交 | set_pte_at / write_cr3 |
lockfile + rename ref | rename 或 CAS metadata_location |
追加新 metadata,再改指针 |
| 回收 | exit_mmap + 页 refcount |
git gc / prune |
expireSnapshots 后再删文件 |
删 LV / 拆共享后再还 PE |
收成一句:底下是可共享的颗粒,上面叠一层或多层索引;换「当前状态」只换根上的指针,不搬叶子。
不同进程的页表可以指向同一张物理页(fork 之后、COW 之前);不同 commit 的 tree 可以指向同一个 blob;新 snapshot 的 manifest 会原样列出没改过的旧 data file;thin 快照的两棵 mapping btree 可以指向同一批数据块。叶子可以继续被旧根引用。索引叠几层,是为了把「找一块」从扫全表收成按路径走,并让共享发生在合适的粒度上。
颗粒怎么命名不一样:页框用 PFN,blob 用内容 hash,data file 用路径,PE 用 PV 上的偏移。颗粒的「不变」程度也不一样。blob 和已提交的 data file 写完就不改;物理页和普通 LV 的 extent 只在被共享、只读时当成这种叶子,独占后可以原地写。Iceberg 这边也不是「不同 manifest 必须指向不同 Parquet」——没改的文件就是被新索引接着指。
这套设计换来的能力可以对照看:
| 能力 | Linux 分页 | Git | Iceberg | LVM |
|---|---|---|---|---|
| 状态切换便宜 | write_cr3,不搬页 |
改 HEAD / checkout | 换 Catalog 里的 metadata 路径 | 改 MDA 指针,不搬 PE |
| 改一点不拷全部 | fork 共享只读页 |
新 commit 复用旧 blob | 新 snapshot 复用旧 data file | thin 快照共享数据块 |
| 读不受写打扰 | 先 set_pte_at,读者走旧映射 |
先写对象再 update-ref |
先写文件再 CAS / rename | 先追加 metadata 再改指针;thin 写新块 |
| 还能回到旧根 | fork 后的共享页是瞬时快照(没有地址空间日志) |
git log 沿 parent |
time travel,旧 snapshot 仍在 | thin 快照;MDA 缓冲里的旧文本 |
| 查找能剪枝 | 按虚拟地址逐级走页表 | 沿路径走 tree | manifest list / 列 bounds,再进 Parquet footer | 按 LE 查到 PE |
| 逻辑名和物理块脱钩 | VA 对 PFN | 文件名在 tree,内容在 blob | 路径在 manifest,行在 data file | LV 偏移对 PE |
| 回收可以往后放 | 页 refcount | git gc / prune |
expireSnapshots 后再删文件 |
删 LV / 拆共享后再还 PE |
用间接层换来廉价快照、隔离写入、共享历史和可剪枝的查找,代价是多一层指针,以及必须另做垃圾回收。
LSM-Tree 也可以用同一副眼镜看:WAL / memtable 是新影子,SST 是不可变颗粒,compaction 是后台重写索引,manifest 是那枚指针。那是另一篇的事。
References
-
Linux 内核
arch/x86/mm/tlb.c—switch_mm_irqs_off()/load_new_mm_cr3()。进程换mm时把next->pgd写入 CR3;注释强调load_cr3()的串行化语义。 ↩ -
Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 3A: System Programming Guide, Part 1(Order 253668),Chapter 4 Paging。§4.5 四级/五级的开启条件;§4.2 从 CR3 起每次切 9 位的迭代走表。全套手册入口:Intel SDM。 ↩ ↩2 ↩3 ↩4
-
Git 手册 git-update-ref — Update the object name stored in a ref safely。源码见
Documentation/git-update-ref.adoc。给定<new-oid> <old-oid>时先验证再写;files backend 用 lockfile +commit_lock_file()完成单条 ref 的原子更新。 ↩ -
Apache Iceberg Table Spec(源码
format/spec.md)— Overview / Optimistic Concurrency / File System Operations。表状态变更写新 metadata,并以原子交换替换旧指针;数据文件写后不可变;Hadoop 表用 rename 提交 metadata。 ↩ ↩2 ↩3 ↩4 -
Apache Iceberg
HadoopTableOperations.java—commit()/renameToFinal()/writeVersionHint()。注释写明 rename 是原子提交;version-hint.text为 best-effort。 ↩ -
Apache Iceberg
BaseMetastoreTableOperations.java中的METADATA_LOCATION_PROP;HMSTablePropertyHelper.java把新路径写入 HMS 表参数。 ↩ -
Linux 内核文档 5-level paging(源码
Documentation/arch/x86/x86_64/5level-paging.rst)— a straight-forward extension of the current page table structure adding one more layer of translation。pgdir_shift默认 39、五级改为 48,见arch/x86/boot/compressed/pgtable_64.c。 ↩ ↩2 -
同上,§4.5.2 Use of CR3 with Ordinary 4-Level Paging and 5-Level Paging;Table 4-12(
CR4.PCIDE = 0)bits 12 及以上为 4K 对齐的 PML4/PML5 物理地址。Linux 写入见arch/x86/include/asm/special_insns.h的native_write_cr3()。 ↩ -
同上,§4.7 Page-Fault Exceptions;Figure 4-12 的 error code(P / W/R / U/S / RSVD)。Linux 对位的注释见
arch/x86/include/asm/trap_pf.h;入口arch/x86/mm/fault.c的exc_page_fault()/do_user_addr_fault()。 ↩ -
Linux 内核
include/linux/pgtable.h—pgd_offset()/pmd_off()。软件页表行走的折叠路径。 ↩ -
Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 3C: System Programming Guide, Part 3(Order 326019-090US,February 2026),§27.4 Guest-State Area、§27.4.1 Guest Register State。VM entry 从这些字段装入处理器状态;guest-state 含 CR0、CR3、CR4。 ↩
-
Volume 3A §4.1.2 Paging-Mode Enabling:VMX transitions allow transitions between paging modes that are not possible using MOV to CR or WRMSR. This is because VMX transitions can load CR0, CR4, and IA32_EFER in one operation. HLAT 用 VMCS 的 HLATP 定位第一张表,见 §4.5.3;细节在 Volume 3C 的 VMCS / VM-execution control。 ↩
-
Prem Kumar Ponuthorai、Jon Loeliger,Version Control with Git 第 3 版(O’Reilly,2022),第 2 章 Foundational Concepts(「Blob Objects and Hashes」「Tree Object and Files」「Commit Objects」)。
hello worldblob3b18e512dba79e4c8300dd08aeb37f8e728b8dad、tree68aba62e560c0ebc3396e8ae9335232cd93a3f60,以及 A blob holds a file’s data but does not contain any metadata about the file or even its name。 ↩ ↩2 ↩3 ↩4 ↩5 -
Tomer Shiran、Jason Hughes、Alex Merced,Apache Iceberg: The Definitive Guide(O’Reilly,2024),第 2 章 The Architecture of Apache Iceberg。data layer 的 data file / delete file;「file format most commonly used is Apache Parquet」;湖上文件按不可变处理。 ↩ ↩2
-
Apache Iceberg spec Manifests — A manifest is an immutable Avro file that lists data files or delete files;这些 manifest 由每个 snapshot 的 manifest list 跟踪。 ↩
-
Apache Iceberg spec Scan Planning(源码 format/spec.md#scan-planning);实现见
ManifestEvaluator.java、InclusiveMetricsEvaluator.java。 ↩ -
Apache Parquet File format;仓库说明
README.md#file-format;ThriftFileMetaData。读者先读 footer,再按 row group / column chunk 定位。 ↩ -
Apache Iceberg 测试夹具
core/src/test/resources/TableMetadataV2Valid.json。current-snapshot-id、parent-snapshot-id、manifest-list的官方样例。 ↩ -
Apache Iceberg
TestReachableFileUtil.java—DataFiles.builder(SPEC).withPath("/path/to/data-a.parquet");路径 API 见DataFiles.java。 ↩ -
Apache Iceberg spec Data File Fields —
file_path(字段 100)、file_format(101)、record_count(103)。manifest 为 Avro,文中 JSON 只用来对照字段。 ↩ -
Linux 内核
mm/memory.c—do_wp_page()。私有映射写过错:能复用则wp_page_reuse(),否则wp_page_copy()。 ↩ -
Linux 内核
mm/memory.c—wp_page_copy()里ptep_clear_flush()之后才set_pte_at(),并说明必须先切换 PTE 再减旧页 mapcount。 ↩ -
Apache Iceberg
SnapshotProducer.java—commit()。apply()出新 snapshot,再ops.commit(base, updated);只对CommitFailedException按表属性重试。 ↩ -
Linux 内核
mm/mmap.c—exit_mmap()。最后一个mm用户离开后 unmap VMA 并free_pgtables()。 ↩ -
Git
builtin/gc.c—prune_packed()。 ↩ -
Apache Iceberg
Snapshot.java;spec Snapshot References;回收见RemoveSnapshots.java。 ↩ -
Red Hat Enterprise Linux 9,Managing LVM volume groups — Extents are the smallest units of space that you can allocate in LVM;默认 4 MiB;The logical extents within the LV correspond one-to-one with physical extents in the VG。 ↩
-
Red Hat Enterprise Linux 7,Appendix E. LVM Volume Group Metadata — A metadata area is a circular buffer. New metadata is appended to the old metadata and then the pointer to the start of it is updated. 元数据为 ASCII,默认在每个 PV 的 metadata area 留一份拷贝。 ↩
-
Linux 内核文档 Thin provisioning(源码
Documentation/admin-guide/device-mapper/thin-provisioning.rst)— 多份虚拟设备共享同一 data volume。实现见drivers/md/dm-thin.c:内部快照克隆 mapping btree 的根节点;写共享块时数据落到新块。 ↩