本节目标:掌握
open的每个参数与pathlib的现代路径操作,理解编码为什么会乱,并学会用原子写保护配置文件。
适用版本:Python 3.12+(实测 3.14.6)
10.2 文件 I/O、pathlib 与编码
上一节 10.1 字符串处理与正则 re
把文本在内存里处理干净了。可程序一重启,内存就没了——要把结果留住,必须落到磁盘。这一节讲文件读写,核心只有一句话:打开文件时,永远显式指定 encoding。
10.2.1 open() 的模式与 encoding
open(path, mode, encoding) 是入口。mode 决定能做什么:
| 模式 | 含义 | 文件不存在时 |
|---|---|---|
r | 只读(默认) | 抛 FileNotFoundError |
w | 写入(清空已有内容) | 创建 |
a | 追加 | 创建 |
x | 独占创建 | 抛 FileExistsError |
b | 二进制(rb/wb) | — |
+ | 读写(r+/w+) | 随基础模式 |
写一个文件再读回来:
from pathlib import Path
p = Path("/tmp/python_book/scratch/ch10/demo.txt")
with open(p, "w", encoding="utf-8") as f:
f.write("第一行\n")
f.write("第二行\n")
f.write("第三行\n")
with open(p, "r", encoding="utf-8") as f:
print(repr(f.read()))
'第一行\n第二行\n第三行\n'
为什么必须显式写 encoding="utf-8"? 因为不写时,open 用的是「平台默认编码」——由 locale.getpreferredencoding(False) 决定。它随环境变化:本机 macOS 是 UTF-8,但很多 Windows 中文环境是 cp936(GBK),Docker 容器里又可能是 POSIX/ASCII。同一份代码在 A 机器上写出的 UTF-8 文件,到 B 机器上按 GBK 读就成了乱码。
import locale, sys
print("preferred:", locale.getpreferredencoding(False))
print("default:", sys.getdefaultencoding())
preferred: UTF-8
default: utf-8
CPython 也把这件事当成问题,3.10 起新增了 EncodingWarning:用 python -X warn_default_encoding 或设 PYTHONWARNDEFAULTENCODING=1,就能让「没写 encoding」的调用报警:
python -X warn_default_encoding -c "open('demo.txt')"
EncodingWarning: 'encoding' argument not specified
另外,二进制模式不能带 encoding:open(path, "rb", encoding="utf-8") 会直接抛 ValueError: binary mode doesn't take an encoding argument。二进制读写用 read_bytes() / write_bytes() 更顺手。
10.2.2 with 与文件对象
open 返回一个文件对象,用完必须关闭,否则缓冲区里的数据可能没落盘、文件句柄会泄漏。7.3 节讲过上下文管理器,with 就是它的典型用法——退出代码块时自动关闭,即使中途抛异常:
with open(p, "r", encoding="utf-8") as f:
print("name:", f.name)
print("mode:", f.mode)
print("encoding:", f.encoding)
print("closed:", f.closed)
print("after with:", f.closed)
name: /tmp/python_book/scratch/ch10/demo.txt
mode: r
encoding: utf-8
closed: False
after with: True
文件对象还带位置指针,f.read() 读完后指针停在末尾,再读就是空串——需要重读时要 f.seek(0)。这个细节在「先读再写」的 r+ 模式里最容易踩坑。
10.2.3 三种读取方式的内存差异
读文本有三种写法:read() 全读成一个大字符串,readlines() 读成行的列表,for line in f 逐行迭代。功能相似,内存占用差一个数量级。用 tracemalloc 量一个 6.9 MB、20 万行的文件:
import tracemalloc
from pathlib import Path
big = Path("big.txt") # 6.9 MB,20 万行
def peak(fn):
tracemalloc.start()
fn()
_, pk = tracemalloc.get_traced_memory()
tracemalloc.stop()
return pk
def with_read():
with open(big, encoding="utf-8") as f:
f.read()
def with_readlines():
with open(big, encoding="utf-8") as f:
f.readlines()
def with_iter():
with open(big, encoding="utf-8") as f:
for _ in f:
pass
print("read() ", round(peak(with_read) / 1e6, 1), "MB")
print("readlines() ", round(peak(with_readlines) / 1e6, 1), "MB")
print("for line in f", round(peak(with_iter) / 1e6, 1), "MB")
read() 13.9 MB
readlines() 16.9 MB
for line in f 0.1 MB
结论很清楚:read() 和 readlines() 会把整个文件拉进内存,逐行迭代几乎不占额外内存。处理日志、大 CSV 这类可能几十上百 MB 的文件,一律用 for line in f。只有文件确实很小、且你需要在内存里整体处理时,才用 read()。
需要更细的控制时,还可以用 f.readline() 读一行、f.read(n) 读 n 个字符,或 f.read(8192) 分块。分块读二进制大文件的标准写法是:
with open(big, "rb") as f:
while chunk := f.read(8192): # 海象运算符
process(chunk)
10.2.4 pathlib.Path 现代路径操作
os.path 的字符串拼接已经过时,pathlib.Path 是 3.4 起的推荐方式。核心是用 / 运算符拼路径,跨平台自动处理分隔符:
from pathlib import Path
base = Path("/tmp/python_book/scratch/ch10")
q = base / "sub" / "nested" / "file.tar.gz"
print(q.name, q.stem, q.suffix, q.suffixes)
print(q.parent)
print(q.parts[-3:])
file.tar.gz file.tar .gz ['.tar', '.gz']
/tmp/python_book/scratch/ch10/sub/nested
('sub', 'nested', 'file.tar.gz')
注意 stem 只去掉最后一个后缀,suffixes 给出全部后缀列表。常用的读写方法直接挂在 Path 上,比 open 更简洁:
t = base / "note.txt"
t.write_text("你好\nworld\n", encoding="utf-8")
print(repr(t.read_text(encoding="utf-8")))
d = base / "out" / "2026" / "09"
d.mkdir(parents=True, exist_ok=True)
print("is_dir:", d.is_dir())
'你好\nworld\n'
is_dir: True
mkdir(parents=True, exist_ok=True) 一次性创建多层目录,且目录已存在时不报错——比 os.makedirs 更清晰。read_text / write_text 一样要显式传 encoding。
10.2.5 glob / rglob / iterdir / walk
遍历目录有四个工具,按用途区分:iterdir() 列当前层全部条目,glob(pattern) 按通配符匹配当前层,rglob(pattern) 递归匹配所有层,Path.walk()(3.12+)像 os.walk 一样自顶向下逐层产出:
print("iterdir:", sorted(x.name for x in (base / "data").iterdir()))
print("glob *.txt:", sorted(x.name for x in (base / "data").glob("*.txt")))
print("rglob *.txt:", sorted(str(x.relative_to(base / "data"))
for x in (base / "data").rglob("*.txt")))
for root, dirs, files in (base / "data").walk():
print("walk:", root.name, sorted(dirs), sorted(files))
iterdir: ['a.txt', 'b.csv', 'sub']
glob *.txt: ['a.txt']
rglob *.txt: ['a.txt', 'sub/c.txt']
walk: data ['sub'] ['a.txt', 'b.csv']
walk: sub [] ['c.txt']
iterdir / glob 返回的是迭代器,遍历一遍就没了,需要多次使用就先 list(...) 固化。Path.walk 是 3.12 才有的,若要兼容更早版本仍得用 os.walk。
10.2.6 file 与脚本定位
脚本里常需要知道「自己在哪里」,好去读同目录的配置或数据文件。__file__ 是当前文件的路径字符串,用 Path 包一层就能定位:
from pathlib import Path
print("__file__:", __file__)
here = Path(__file__).parent
print("parent:", here)
print("config:", here / "config.json")
__file__: /private/tmp/python_book/scratch/ch10/t102_encoding.py
parent: /private/tmp/python_book/scratch/ch10
config: /private/tmp/python_book/scratch/ch10/config.json
(macOS 上 /tmp 是指向 /private/tmp 的符号链接,所以 __file__ 显示成解析后的路径,这是正常的。)
不要用 os.getcwd() 或相对路径定位同目录文件——当前工作目录随启动方式变化,从别的目录运行脚本就会找不到文件。Path(__file__).parent 则始终指向脚本所在目录,这是可靠的做法。需要绝对路径时再加 .resolve()。
10.2.7 编码排查:BOM / utf-8-sig / GBK
乱码几乎都源于「写入时的编码」与「读取时的编码」不一致。最常见的是 UTF-8 BOM:Windows 记事本、Excel 导出的文件常在开头写三个字节 EF BB BF,用普通 utf-8 读会多出一个不可见的 :
bom_file = base / "bom.txt"
bom_file.write_text("hello", encoding="utf-8-sig") # 写时带 BOM
print(bom_file.read_bytes())
print(repr(bom_file.read_text(encoding="utf-8")))
print(repr(bom_file.read_text(encoding="utf-8-sig")))
b'\xef\xbb\xbfhello'
'\ufeffhello'
'hello'
用 utf-8-sig 读会自动吞掉 BOM,这是处理外部 CSV/文本的稳妥选择(没有 BOM 时它也正常工作)。第二种常见情况是 GBK 文件被当 UTF-8 读:
gbk_file = base / "gbk.txt"
gbk_file.write_bytes("中文测试\n".encode("gbk"))
try:
gbk_file.read_text(encoding="utf-8")
except UnicodeDecodeError as e:
print("utf-8 fails:", e.reason)
print("gbk:", repr(gbk_file.read_text(encoding="gbk")))
utf-8 fails: invalid continuation byte
gbk: '中文测试\n'
排查顺序是:先 read_bytes() 看原始字节,再用 errors="replace" 试解码,看哪些字节被替换成 �,就能锁定问题。记住三个数字:UTF-8 里一个汉字 3 字节,GBK 里一个汉字 2 字节——字节数对不上,就是编码猜错了。
10.2.8 临时文件与原子写
临时文件用 tempfile 创建,避免手工拼 /tmp 带来的竞态与权限问题:
import tempfile
from pathlib import Path
with tempfile.TemporaryDirectory() as d:
f = Path(d) / "x.txt"
f.write_text("temp", encoding="utf-8")
print("inside:", f.read_text(encoding="utf-8"))
print("cleaned:", not Path(d).exists())
inside: temp
cleaned: True
原子写解决的是「写到一半程序崩溃,配置文件被截断」的问题。做法是:先写到同目录的临时文件,flush 后再用 os.replace 一次性改名覆盖目标。os.replace 在同一文件系统内是原子操作——要么旧文件完整,要么新文件完整,不会出现半截文件:
import os, tempfile
from pathlib import Path
target = base / "config.json"
with tempfile.NamedTemporaryFile(
"w", encoding="utf-8", dir=target.parent, delete=False, suffix=".tmp"
) as tf:
tf.write('{"v": 2}\n')
tmp_name = tf.name
os.replace(tmp_name, target)
print(target.read_text(encoding="utf-8").strip())
{"v": 2}
两个要点:临时文件必须和目标在同一个目录(dir=target.parent),否则跨文件系统改名就不是原子的;delete=False 是为了让 with 退出后文件仍在,由 os.replace 接管。写配置、写数据快照、写缓存时都该用这个模式。
更完整的文件与序列化话题见专题 Python 文件 IO 与数据序列化 。
小结
open的模式r/w/a/x/b/+各司其职;w会清空文件,别拿它做追加。- 永远显式指定
encoding="utf-8":默认编码随平台变化,是跨平台乱码的根源;二进制模式不能带encoding。 with自动关闭文件;read()/readlines()全量入内存,大文件一律用for line in f逐行迭代。pathlib.Path用/拼路径,read_text/write_text/mkdir(parents=True)比os.path清爽;glob/rglob/iterdir/walk覆盖遍历需求。- 定位同目录文件用
Path(__file__).parent,不要用当前工作目录。 - UTF-8 BOM 用
utf-8-sig处理;GBK 乱码先看原始字节;写文件用tempfile+os.replace保证原子性。
文本落了盘,接下来要面对结构化的数据格式:配置是 TOML,接口数据是 JSON,表格导出是 CSV。下一节 10.3 JSON / CSV / TOML 与序列化安全
会逐个拆解,并讲清 pickle 为什么不能碰不可信数据。
阅读导航:上一节:10.1 字符串处理与正则 re · 下一节:10.3 JSON / CSV / TOML 与序列化安全 。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。