当前位置:首页>python>Python 高级容器:collections`模块完全指南

Python 高级容器:collections`模块完全指南

  • 2026-10-11 06:26:39
Python 高级容器:collections`模块完全指南

Python 内置的 dict、list、set、tuple 已经够用,但在特定场景下,标准库 collections 提供了更顺手、更高效的选择。本文基于 Python 3.12,系统梳理 collections 中六大核心容器:namedtuple、deque、Counter、defaultdict、OrderedDict、ChainMap。


一、模块概览

collections 实现了面向特定场景的容器类型,可作为内置容器的替代或增强:

类型
用途
namedtuple()
带字段名的元组工厂函数
deque
双端队列,两端增删均为 O(1)
Counter
可哈希对象计数器
defaultdict
缺失键时自动调用工厂函数
OrderedDict
有序字典,擅长重排序操作
ChainMap
多字典合并为单一视图
import collectionsprint([name for name indir(collections) ifnot name.startswith("_")])# ['ChainMap', 'Counter', 'OrderedDict', 'UserDict', 'UserList',#  'UserString', 'defaultdict', 'deque', 'namedtuple']

二、namedtuple 具名元组

2.1 为什么需要它

普通元组靠索引访问,可读性差。namedtuple 给每个位置赋予名字,既能按字段名访问,又保持元组的不可变性。

from collections import namedtuplePoint = namedtuple("Point", ["x", "y"])p = Point(11, y=22)print(p.x + p.y)   # 33print(p[0] + p[1]) # 33x, y = p            # 支持解包print(p)            # Point(x=11, y=22)

2.2 创建语法

collections.namedtuple(    typename,           # 新子类的类名    field_names,        # 字段名,如 ['x', 'y'] 或 'x y'    *,    rename=False,       # 无效字段名自动重命名    defaults=None,      # 默认值元组    module=None,        # 设置 __module__ 属性)

2.3 特有方法与属性

方法/属性
说明
_make(iterable)
从序列创建实例
_asdict()
转为字典(3.8+ 返回普通 dict)
_replace(**kwargs)
返回替换部分字段的新实例
_fields
字段名元组
_field_defaults
带默认值的字段字典
Point = namedtuple("Point", ["x", "y"])p = Point(11, 22)p2 = Point._make([33, 44])print(p2._asdict())          # {'x': 33, 'y': 44}  ← 3.8+ 为 dict,非 OrderedDictp3 = p2._replace(x=99)print(p3)                    # Point(x=99, y=44)# 组合字段创建新类型Color = namedtuple("Color", "red green blue")Pixel = namedtuple("Pixel", Point._fields + Color._fields)pi = Pixel(11, 22, 128, 255, 0)print(pi)  # Pixel(x=11, y=22, red=128, green=255, blue=0)# 默认值(3.7+)Account = namedtuple("Account", ["type", "balance"], defaults=[0])print(Account._field_defaults)  # {'balance': 0}a = Account("premium")print(a)  # Account(type='premium', balance=0)

2.4 Python 3.12 下的替代方案

新项目更推荐 dataclass(frozen=True) 或 typing.NamedTuple,IDE 支持更好、类型提示更完整:

from dataclasses import dataclass@dataclass(frozen=True)classPoint:    x: int    y: intp = Point(11, 22)print(p.x)  # 11

namedtuple 仍适合轻量、无依赖的场景,以及需要与旧代码兼容的项目。


三、deque 双端队列

3.1 核心特性

deque(double-ended queue,读作 "deck")支持两端 O(1) 的追加与弹出,线程安全,内存效率高。

相比 list,deque 在 pop(0)、insert(0, v) 等操作上不会触发 O(n) 的整体移动。

from collections import dequed = deque("ghi")d.append("j")        # 右端追加d.appendleft("f")    # 左端追加print(d)             # deque(['f', 'g', 'h', 'i', 'j'])print(d.pop())       # jprint(d.popleft())   # fprint(list(d))       # ['g', 'h', 'i']

3.2 主要方法

方法
说明
append(x)
 / appendleft(x)
右/左端追加
extend(iterable)
 / extendleft(iterable)
右/左端批量扩展
pop()
 / popleft()
右/左端弹出
rotate(n=1)
向右循环移动 n 步(负数则向左)
maxlen
只读属性,最大长度限制
clear()
 / copy() / count(x) / index(x)
常规操作

3.3 定长队列(滑动窗口)

设置 maxlen 后,超出长度时从另一端自动弹出旧元素,适合"最近 N 条记录"场景:

recent = deque(maxlen=3)for item inrange(5):    recent.append(item)print(list(recent))# [0]# [0, 1]# [0, 1, 2]# [1, 2, 3]# [2, 3, 4]

3.4 实战:轮询调度器

defroundrobin(*iterables):"""roundrobin('ABC', 'D', 'EF') --> A D E B F C"""    iterators = deque(map(iter, iterables))while iterators:try:whileTrue:yieldnext(iterators[0])                iterators.rotate(-1)except StopIteration:            iterators.popleft()

四、Counter 计数器

4.1 基本用法

Counter 是 dict 的子类,用于统计可哈希对象出现次数。访问不存在的键返回 0,不会抛出 KeyError。

from collections import Counter# 从列表计数L = ["red", "blue", "red", "green", "blue", "blue"]print(Counter(L))  # Counter({'blue': 3, 'red': 2, 'green': 1})# 从字符串计数print(Counter("gallahad"))  # Counter({'g': 1, 'a': 3, 'l': 2, 'h': 1, 'd': 1})# 从字典初始化print(Counter({"red": 4, "blue": 2}))  # Counter({'red': 4, 'blue': 2})

4.2 特有方法

c = Counter(a=4, b=2, c=0, d=-2)# 按首次出现顺序,重复计数值次print(list(c.elements()))  # ['a', 'a', 'a', 'a', 'b', 'b']# 最常见的 n 个元素print(Counter("abracadabra").most_common(3))# [('a', 5), ('b', 2), ('r', 2)]# 减法(不是替换,而是减去)c.subtract({"a": 1, "b": 2, "c": 3, "d": 4})print(c)  # Counter({'a': 3, 'b': 0, 'c': -3, 'd': -6})# 总计数(3.10+)print(Counter(a=10, b=5, c=3).total())  # 18

4.3 数学运算(3.10+)

c = Counter(a=3, b=-1, c=-2)d = Counter(a=6, b=2, c=-4)print(c + d)  # Counter({'a': 9, 'b': 1})print(c - d)  # Counter({'c': 2})  ← 结果 ≤ 0 的键被删除print(c & d)  # Counter({'a': 3})  ← 取各键最小值print(c | d)  # Counter({'a': 6, 'b': 2, 'c': -2})  ← 取各键最大值# 一元运算print(+c)  # Counter({'a': 3, 'b': -1, 'c': -2})print(-c)  # Counter({'b': 1, 'c': 2})# 富比较(3.10+:不存在的元素视为计数 0)print(c == d)   # Falseprint(c >= d)   # False

注意:Counter.update() 是累加计数,不是字典式的覆盖更新;也没有 fromkeys() 方法。


五、defaultdict — 带默认值的字典

5.1 基本用法

defaultdict 在键不存在时,自动调用 default_factory 生成默认值,而不是抛出 KeyError。

from collections import defaultdictd = defaultdict(int)print(d["a"])  # 0(自动调用 int())

5.2 __missing__ 机制

只有 d[key] 会触发 __missing__;d.get(key) 仍返回 None,行为与普通 dict 一致。

5.3 常见模式

分组聚合(list 工厂)

s = [("yellow", 1), ("blue", 2), ("yellow", 3), ("blue", 4), ("red", 1)]d = defaultdict(list)for k, v in s:    d[k].append(v)print(dict(d))# {'yellow': [1, 3], 'blue': [2, 4], 'red': [1]}

字符计数(int 工厂)

d = defaultdict(int)for ch in"mississippi":    d[ch] += 1print(sorted(d.items()))  # [('i', 4), ('m', 1), ('p', 2), ('s', 4)]

去重集合(set 工厂)

d = defaultdict(set)for color, val in [("red", 1), ("blue", 2), ("red", 3), ("blue", 4), ("red", 1)]:    d[color].add(val)print(dict(d))  # {'red': {1, 3}, 'blue': {2, 4}}

自定义默认值(lambda 工厂)

d = defaultdict(lambda: "<missing>")d.update(name="John", action="ran")print("%(name)s %(action)s to %(object)s" % d)# John ran to <missing>

六、OrderedDict 有序字典

6.1 还需要它吗?

自 Python 3.7 起,内置 dict 已保证插入顺序;3.8 起 dict 也支持 reversed()。因此,单纯为了"记住顺序",dict 通常已足够。

OrderedDict 仍有独特价值:

特性
普通 dict
OrderedDict
插入顺序
✅(3.7+ 保证)
✅
move_to_end()
❌
✅
popitem(last=False)
 弹出首项
❌
✅
顺序敏感的相等比较
❌
✅
频繁重排序性能
一般
更优(LRU 缓存等)
from collections import OrderedDict# 顺序敏感比较od1 = OrderedDict([(1, 1), (2, 2)])od2 = OrderedDict([(2, 2), (1, 1)])print(od1 == od2)  # Falsed1 = {1: 1, 2: 2}d2 = {2: 2, 1: 1}print(d1 == d2)  # True

6.2 常用方法

d = OrderedDict.fromkeys("abcde")# popitem:LIFO(默认)或 FIFOprint(d.popitem())       # ('e', None)print(d.popitem(last=False))  # ('a', None)# move_to_end:将元素移到末尾或开头d = OrderedDict.fromkeys("abcde")d.move_to_end("b")print(d)  # OrderedDict([('a', None), ('c', None), ('d', None), ('e', None), ('b', None)])d.move_to_end("b", last=False)print(d)  # OrderedDict([('b', None), ('a', None), ('c', None), ('d', None), ('e', None)])# 逆序迭代print(list(reversed(d)))  # ['e', 'd', 'c', 'a', 'b']

6.3 Python 3.12 变化:repr 格式更新

3.12 起,OrderedDict 的 repr 输出改为与普通 dict 一致的花括号格式:

from collections import OrderedDictod = OrderedDict([("c", 1), ("b", 2), ("a", 3)])# Python 3.12+print(repr(od))# OrderedDict({'c': 1, 'b': 2, 'a': 3})# Python 3.11 及更早# OrderedDict([('c', 1), ('b', 2), ('a', 3)])

若测试或日志中硬编码了旧格式字符串,升级到 3.12 后需要相应调整。

6.4 实战:简易 LRU 缓存

from collections import OrderedDictclassLRUCache:def__init__(self, capacity: int):self.cache = OrderedDict()self.capacity = capacitydefget(self, key):if key notinself.cache:return -1self.cache.move_to_end(key)returnself.cache[key]defput(self, key, value):if key inself.cache:self.cache.move_to_end(key)self.cache[key] = valueiflen(self.cache) > self.capacity:self.cache.popitem(last=False)  # 弹出最久未使用的

七、ChainMap — 链式字典

7.1 基本用法

ChainMap 将多个字典组合为单一视图,查找时按顺序从前到后搜索,写入只作用于第一个字典。

from collections import ChainMapbaseline = {"music": "bach", "art": "rembrandt"}adjustments = {"art": "van gogh", "opera": "carmen"}cm = ChainMap(adjustments, baseline)print(cm["art"])    # van gogh(adjustments 优先)print(cm["music"])  # bach(baseline 中找到)print(list(cm))     # ['music', 'art', 'opera']

7.2 特有属性与方法

print(cm.maps)# [{'art': 'van gogh', 'opera': 'carmen'}, {'music': 'bach', 'art': 'rembrandt'}]# 创建子上下文(不影响父级)child = cm.new_child({"new_key": 666})print(child)# ChainMap({'new_key': 666}, {'art': 'van gogh', ...}, {'music': 'bach', ...})# 跳过第一个映射print(cm.parents)# ChainMap({'music': 'bach', 'art': 'rembrandt'})

7.3 实战场景

命令行参数 > 环境变量 > 默认值

import osimport argparsefrom collections import ChainMapdefaults = {"color": "red", "user": "guest"}parser = argparse.ArgumentParser()parser.add_argument("-u", "--user")parser.add_argument("-c", "--color")namespace = parser.parse_args()cli_args = {k: v for k, v invars(namespace).items() if v isnotNone}config = ChainMap(cli_args, os.environ, defaults)print(config["color"])  # 命令行 > 环境变量 > 默认值print(config["user"])

模拟 Python 作用域链

import builtinsfrom collections import ChainMappylookup = ChainMap(locals(), globals(), vars(builtins))

多模块共享库存(引用存储,修改实时反映)

from collections import ChainMaptoys = {"Blocks": 30, "Monopoly": 20}computers = {"iMac": 1000, "Chromebook": 1000}clothing = {"Jeans": 40, "T-shirt": 10}inventory = ChainMap(toys, computers, clothing)print(inventory["Monopoly"])  # 20toys["Nintendo"] = 20# 底层字典修改会反映到 ChainMapprint(inventory["Nintendo"])  # 20

3.9+ 支持 | 和 |= 合并运算符(PEP 584)。


八、选型速查表

需求
推荐容器
固定结构、不可变记录
namedtuple
 / dataclass(frozen=True)
双端队列、滑动窗口、BFS
deque
词频统计、Top-K
Counter
分组聚合、自动初始化
defaultdict
LRU 缓存、顺序敏感比较
OrderedDict
多层配置合并、作用域模拟
ChainMap
仅需插入顺序
普通 dict(3.7+)即可

参考

  • Python 3.12 官方文档 — collections [1]

  • PEP 584 — Merge operators for dict [2]

  • gh-101446: OrderedDict repr 变更 [3]

引用链接

[1] Python 3.12 官方文档 — collections: https://docs.python.org/zh-cn/3.12/library/collections.html[2] PEP 584 — Merge operators for dict: https://peps.python.org/pep-0584/[3] gh-101446: OrderedDict repr 变更: https://github.com/python/cpython/issues/101446

最新文章

随机文章