当前位置:首页>python>Python学习 | 06 Pythonic Iteration:Comprehension与更自然的循环写法

Python学习 | 06 Pythonic Iteration:Comprehension与更自然的循环写法

  • 2026-10-10 19:46:29
Python学习 | 06 Pythonic Iteration:Comprehension与更自然的循环写法

前五课,我们已经能够写出真正的 sequence-processing loop:

seq = ”ATGCGT”k = 3kmers = []for i in range(len(seq) - k + 1):    kmer = seq[i:i + k]    kmers.append(kmer)print(kmers)

这段代码完全正确,而且非常重要。

现在 Lesson 06 要做的不是推翻它,而是认识 Python 中一种极其常见的模式:

建立空容器↓for 遍历↓计算一个结果↓append 到容器

当逻辑足够简单时,Python 通常会把它写成 comprehension。

例如上面的 k-mer:

kmers = [    seq[i:i + k]    for i in range(len(seq) - k + 1)]kmers

这就是这一课的核心。

不过有一条原则要从一开始就明确:

Pythonic 不等于“代码越短越好”。

我们真正关心的是:

哪种写法最直接地表达这段代码想做什么。

1. List comprehension:把“循环 + 收集结果”写在一起

1.1 从普通 for 到 list comprehension

先看普通写法:

reads = [    ”ATGCGT”,    ”ATCC”,    ”GGGAAA”]lengths = []for read in reads:    lengths.append(len(read))print(lengths)

这里的逻辑其实只有一句话:

对每条 read 计算 len(read),把结果放进一个 list。

所以 Python 可以写:

lengths = [len(read) for read in reads]print(lengths)

基本结构是:

[    expression    for item in iterable]

也就是:

对 iterable 中的每个 item↓计算 expression↓收集到新的 list

假如想把reads全部变成大写:

reads = [”atgcgt”, ”atcc”, ”gggaaa”]

普通 for:

upper_reads = []for read in reads:    upper_reads.append(read.upper())print(upper_reads)

comprehension:

upper_reads = [read.upper() for read in reads]print(upper_reads)

所以你可以先把 list comprehension 理解成:

for + append

的一种紧凑表达。

1.2 和 R 的思维对照

对你来说,这个概念其实不会陌生。

R 中:

reads <- c(”ATGC”, ”GGAA”, ”ATCG”)nchar(reads)

可以直接得到每条 sequence 的长度,因为很多 R 函数本身就是 vectorized 的。

Python list:

reads = [”ATGC”, ”GGAA”, ”ATCG”]

不是 R vector,所以不能假定:

len(reads)

因为它问的是:

这个 list 有多少个元素?

要计算每个元素的长度:

[len(read) for read in reads]

2. 在 comprehension 中筛选元素

2.1 for ... if ...:只保留满足条件的对象

在 comprehension 末尾加入 if,可以只收集满足条件的元素。

reads = [    ”ATGCGT”,    ”ATGC”,    ”GGCCAATT”,    ”AT”,    ”CCGTA”]

只保留长度至少为 5 的 reads。

普通写法:

long_reads = []for read in reads:    if len(read) >= 5:        long_reads.append(read)print(long_reads)

comprehension:

long_reads = [    read    for read in reads    if len(read) >= 5]long_reads

结构是:

[    expression    for item in iterable    if condition]

阅读时不要从左到右机械翻译。

可以按照这个顺序理解:

for read in reads        ↓if len(read) >= 5        ↓保留 read

再比如过滤掉含 N 的 reads:

reads = [    ”ATGCGT”,    ”ATGCNT”,    ”GGCCAA”,    ”NNNNNN”]
valid_reads = [    read    for read in reads    if ”N” not in read]print(valid_reads)

2.2 筛选和转换可以同时进行

例如:

只处理没有 N 的 read,然后计算长度。

length = [    len(read)    for read in reads    if ”N” not in read]print(length)

3. 两种 if 不要混淆:筛选 vs if/else 转换

这是 list comprehension 最容易写乱的地方之一。

3.1 只有 if:表示筛选

numbers = [1, 2, 3, 4, 5]

只保留偶数:

even_numbers = [    x    for x in numbers    if x % 2 == 0]print(even_numbers)

所以:

... for x in data if condition

表达的是:

filter

3.2 if ... else ...:每个元素都保留,但决定产生什么结果

例如给 read 打标签:

reads = [    ”ATGCGT”,    ”ATGC”,    ”GGCCAATT”]

如果长度至少为 5:pass,否则 short

labels = [    ”pass” if len(read) >=5 else ”short”    for read in reads]labels

筛选:

[    read    for read in reads    if len(read) >= 5]

而 conditional expression:

[    ”pass” if len(read) >=5 else ”short”    for read in reads]

因为:

”pass” if condition else ”short”

本身就是一个完整的 expression。

类似:

如果条件成立,expression 的结果是 ”pass”否则结果是 ”short”

4. 把 comprehension 用到 sequence algorithm

这一部分比拿 [x * 2 for x in numbers] 举例重要得多。

4.1 一行生成所有 k-mer

Lesson 05 我们写过:

seq = ”ATGCGT”k = 3kmers = []for i in range(len(seq) - k + 1):    kmer = seq[i:i+k]    kmers.append(kmer)print(kmers)

List comprehension 可以直接把每一个滑动窗口收集为 k-mer。

kmers = [    seq[i:i+k]    for i in range(len(seq) - k + 1)]print(kmers)

甚至可以写成一行:

kmers = [seq[i:i+k] for i in range(len(seq) - k + 1)]print(kmers)

4.2 找出所有 motif positions

Lesson 05:

seq = ”GATATATGCATATACTT”motif = ”ATAT”positions = []for i in range(len(seq) - len(motif) + 1):    if seq[i:i + len(motif)] == motif:        positions.append(i + 1)positions

Comprehension 可以收集所有满足 motif 匹配条件的位置。

positions = [    i + 1    for i in range(len(seq) - len(motif) + 1)    if seq[i:i+len(motif)] == motif]positions

4.3 什么时候 range() 仍然是正确选择?

上一课我们说:

for i, base in enumerate(seq):

通常比:

for i in range(len(seq)):    base = seq[i]

更自然。

但不要因此形成:

range(len(...)) 永远是不 Pythonic 的。

比如 k-mer:

seq[i:i + k]

我们真正需要的是:

每一个合法的 window start index

所以:

range(len(seq) - k + 1)

就是非常自然的写法。

我真正需要的是元素,还是位置?

如果只需要元素:

for read in reads:

需要:

index + element

使用:

enumerate()

需要:

一系列 window start positions

使用

range()

这才是 Pythonic iteration 真正应该建立的判断能力。

5. Dictionary 和 Set comprehension

Comprehension 不只可以创建 list。

Python 也提供:

list comprehensiondict comprehensionset comprehension

5.1 Dictionary comprehension:生成 key: value

例如:

sequences = {    ”seq1”: ”ATGCGT”,    ”seq2”: ”ATCC”,    ”seq3”: ”GGGAAA”}

我们想建立:

sequence ID → sequence length

普通写法:

lengths = {}for seq_id, seq in sequences.items():    lengths[seq_id] = len(seq)print(lengths)

Dictionary comprehension:

lengths = {    seq_id: len(seq)    for seq_id, seq in sequences.items()}print(lengths)

基本结构:

{    key_expression: value_expression    for item in iterable}

和 list comprehension 最大的区别只是:

list→ 收集 expressiondict→ 收集 key: value

还可以加入筛选。

例如:

只保存长度至少为 5 的 sequence。

long_sequences = {    seq_id: seq    for seq_id, seq in sequences.items()    if len(seq) >= 5}print(long_sequences)

5.2 Set comprehension:自动得到 unique elements

例如:

seq = ”ATGCGT”k = 2

生成所有 unique 2-mers:

unique_kmers = {    seq[i:i+k]    for i in range(len(seq) - k +1)}print(unique_kmers)

Set comprehension: {read for read in reads}

dictionary comprehension: {seq_id: seq for seq_id, seq in sequences.items()}

kmers = [”ATG”, ”TGC”, ”ATG”, ”GCG”, ”TGC”]unique_kmers = {kmer for kmer in kmers}print(unique_kmers)
不过如果你只是为了去重,set更加直接
set(kmers)

6. enumerate()、zip() 与 comprehension 怎么配合?

Lesson 05 已经认识了这两个工具。这一课主要解决“什么时候选哪个”。

6.1 enumerate():需要“位置 + 元素”

例如:

seq = ”ATGC”

想得到:

[    (0, ”A”),    (1, ”T”),    (2, ”G”),    (3, ”C”)]

可以:

positions = [    (i, base)    for i, base in enumerate(seq)]print(positions)

如果只想找到所有 "G" 的位置:

seq = ”ATGGC”g_positions = [    i    for i, base in enumerate(seq)    if base == ”G”]print(g_positions)
seq.find(”G”)

6.2 zip():需要“多个 iterable 按位置配对”

例如:

seq_ids = [”seq1”, ”seq2”, ”seq3”]sequences = [”ATGC”, ”GGAA”, ”ATCGT”]

可以建立 dictionary:

seq_dict = {    seq_id: seq    for seq_id, seq in zip(seq_ids, sequences)}print(seq_dict)

当然这里其实还有一个更直接的写法:

seq_dict = dict(zip(seq_ids, sequences))print(seq_dict)

可以先建立这张表:

你真正需要什么
更自然的工具
只需要元素
for item in data
元素 + index
enumerate(data)
多个 iterable 同时遍历
zip(a, b)
自己生成整数范围
range(...)
sequence window 的起始位置
range(...)

7. Comprehension 不等于 vectorization

这一点我想专门强调,因为对 R 用户非常重要。

你可能看到

lengths = [len(read) for read in reads]

觉得:

这是不是 Python 的 vectorized operation?

不是。

List comprehension 本质上仍然是在逐个迭代元素。

概念上它仍然类似:

lengths = []for read in reads:    lengths.append(len(read))

只是 Python 提供了一种更适合表达:

遍历 → 转换 → 收集

的语法。

真正到了 NumPy,我们才会看到另一种完全不同的思维:

import numpy as npx = np.array([1, 2, 3])x * 2

这里才更加接近你熟悉的 R vectorization。

因此目前最好区分:

Python list comprehension→ iterationNumPy array operation→ vectorization

8. 什么时候应该用 comprehension,什么时候继续写普通 for?

这是这一课真正比“会写语法”更重要的地方。

适合 comprehension:一个简单的转换

例如:

lengths = [len(read) for read in reads]

非常清楚:

每条 read → length

valid_reads = [    read    for read in reads    if ”N” not in read]

非常自然:

从 reads 中选出不含 N 的。

lengths = [    len(read)    for read in reads    if ”N” not in read]

但是下面这些情况,通常继续写普通 for 更好。

例如:

for read in reads:    gc_count = read.count(”G”) + read.count(”C”)    gc_content = gc_count / len(read)    if gc_content > 0.6:        print(read, gc_content)

硬压成一行不会更 Pythonic,只会更难读。


例如 nucleotide counting:

counts = {    ”A”: 0,    ”C”: 0,    ”G”: 0,    ”T”: 0}for base in seq:    counts[base] += 1

这就是很好的普通 loop。

不应该为了“所有循环都改成 comprehension”而强行改写。

后面我们会学:

Counter(seq)

那是因为已经有一个更合适的数据结构工具,不是因为 for 不好。


例如:

for read in reads:    print(read)

不要写:

[print(read) for read in reads]

虽然 Python 可以执行,但这不是好的用法。

为什么?

因为 comprehension 的语义应该是:

创建一个新的 collection。

而这里我们的真实目的只是:print

并不想创建 list。

普通 for 反而最清楚。


例如:

for base in seq:    if base == ”N”:        print(”Invalid sequence”)        break

普通 for 就非常自然。

不要为了省几行代码把控制逻辑塞进复杂 expression。


所以最终可以记成:

简单的transform / filter / collect        ↓comprehension复杂的logic / state / side effect / control flow        ↓普通 for

9. 一个综合例子:从 reads 到 QC 后的 k-mers

现在把前六课连起来。

给定:

reads = [    ”ATGCGT”,    ”ATGCNT”,    ”GGCCAA”,    ”AT”,    ”ATGCGT”]

我们的第一步是:

去掉含 N 的 reads,并且只保留长度至少为 3 的 reads。

clean_reads = [    read    for read in reads    if ”N” not in read and len(read) >= 3]print(clean_reads)

如果只是想得到 unique reads:

unique_reads = set(clean_reads)unique_reads

然后,对于一条 sequence:

seq = ”ATGCGT”k=3kmers = [    seq[i:i+k]    for i in range(len(seq) - k + 1)]kmers

再得到 unique k-mers:

unique_kmers = set(kmers)unique_kmers

10. 本课练习

这一课我仍然只留几道真正值得写的。

Exercise A|普通 for → comprehension

给定:

reads = [    ”ATGCGT”,    ”ATCC”,    ”GGGAAA”,    ”A”]

先用普通 for 得到每条 read 的长度,然后改写为list comprehension

read_len = []for read in reads:    read_len.append(len(read))read_len
read_len = [    len(read)    for read in reads]read_len

Exercise B|筛选 reads

仍然使用:

reads = [    ”ATGCGT”,    ”ATGCNT”,    ”GGCCAA”,    ”NNNNNN”,    ”AT”]

用一个 list comprehension,只保留:

不含 N 并且 长度 >= 5

clean_reads = [    read    for read in reads    if ”N” not in read and len(read) >= 5]clean_reads

Exercise C|if 筛选和 if/else 转换

给定:

reads = [    ”ATGCGT”,    ”AT”,    ”GGCCAA”]

分别写两个 comprehension。

第一个:

只保留长度 >= 5 的 read

第二个:

每条 read 都保留,长度 >= 5 → "pass" 否则 → "short"

然后解释:

为什么两个 comprehension 中 if 的位置不同?

[    read    for read in reads    if len(read) >= 5]
[    ”pass” if len(read) >= 5 else ”short”    for read in reads]

因为第一个是筛选read,第二个只是定义label,没有对read进行筛选。并且"pass" if len(read) >= 5 else "short"本身就是一个expression,

Exercise D|k-mer composition

给定:

seq = ”CAATCCAAC”k = 5

使用 list comprehension 生成所有 5-mer。

不要先运行。

先回答:

sequence length = ? k = 5

应该产生多少个 k-mer?

使用:

n - k + 1

自己预测结果数量。

然后再写代码。

len(seq) - k + 1
[    seq[i:i+k]    for i in range(len(seq) - k + 1)]

Exercise E|Motif positions

给定:

seq = ”GATATATGCATATACTT”motif = ”ATAT”

使用 list comprehension 找到所有:

1-based motif positions

要求最终得到一个:

positions = [...]

而不是直接 print()。

positions = [    i + 1    for i in range(len(seq) - len(motif) + 1)    if seq[i:i+len(motif)] == motif]positions

Exercise F|Dictionary comprehension

给定:

sequences = {    ”seq1”: ”ATGCGT”,    ”seq2”: ”ATCC”,    ”seq3”: ”GGGAAA”}

建立:

lengths

让结果表示:

seq_id → sequence length

然后进一步只保留:

length >= 5

的sequence。

seq_len = {    seq_id: len(seq)    for seq_id, seq in sequences.items()}seq_len
long_seq = {    seq_id:seq    for seq_id, seq in sequences.items()    if len(seq) >= 5}long_seq

最新文章

随机文章