当前位置:首页>python>一天一个科研小技巧——Python复刻多维气泡相关性散点图

一天一个科研小技巧——Python复刻多维气泡相关性散点图

  • 2026-10-11 05:35:17
一天一个科研小技巧——Python复刻多维气泡相关性散点图

审稿人爱看怎样的图表?🤔 往往不是越复杂越花哨的3D图🚫,而是能用最简明的二维平面📐,讲出多维度数据故事的散点图。📊✨

今天,我们把目光锁定在顶刊经常翻牌子的👑多维气泡相关性散点图(Bubble Scatter Plot)上。🔮 用一段 Python 代码🐍,教你如何把一张平平无奇的散点图,升级为包含“趋势、分组、权重、统计检验”四合一的信息载体。🚀

🎨高颜值气泡图:一图四用的黄金组合

图中的散点图其实是一个“信息组合拳”💪——它把相关趋势 + 分组映射 + 权重映射 + 统计检验放在了一起。

X/Y轴坐标:展示两个核心变量(例如连锁店密度的年均变化与肥胖率的年均变化)的变化趋势。

气泡大小(Size):映射了第三个维度的信息(如国家人口规模、总体样本量或总销售额),拒绝被小样本国家掩盖的大趋势。

颜色分类(Color):直观区分不同组别(图中按世界银行收入水平分为 LMICs、HICs、UMICs)。

相关系数(r 值):直接在图内标注出 Pearson 相关系数,统计结果一目了然。

顶刊论文常用它来同时对比多个维度,例如流行病学中环境因子变化与疾病负担关联,或社会经济学中的宏观指标分析。🌍 下面我们用 Python 的matplotlib把它画出来。👀

(注:数据运用随机数替代,实际操作中替换为你的数据即可)

🐍核心代码思路

import numpy as npimport matplotlib.pyplot as pltfrom scipy.stats import pearsonrfrom matplotlib.lines import Line2D# ================= 1. 配色与数据模拟 =================colors_map = {    'LMICs': {'face': '#C585B3', 'edge': '#A0608F'},    'HICs':  {'face': '#81A9C6', 'edge': '#5A84A4'},    'UMICs': {'face': '#D8A492', 'edge': '#B87B66'}}def generate_group_data(seed_offset):    np.random.seed(42 + seed_offset)    # 模拟相关性方向     slope_sign = -1 if seed_offset == 2 else 1    x1 = np.random.normal(8, 4, 15)    y1 = np.random.normal(slope_sign * 0.4 * x1 + 3, 2, 15)    x2 = np.random.normal(0, 2, 20)    y2 = np.random.normal(slope_sign * 0.2 * x2 + 1, 1.5, 20)    x3 = np.random.normal(-3, 3, 18)    y3 = np.random.normal(slope_sign * 0.3 * x3 + 4, 1.8, 18)    # 气泡大小跨度拉大,呈现明显的权重差异感    sizes = np.random.uniform(40, 900, 15 + 20 + 18)    return (x1, x2, x3), (y1, y2, y3), sizes# ================= 2. 核心绘图函数 =================def draw_bubble_plot(ax, df_x, df_y, sizes, title_str, x_label, y_label=None):    x_lm, x_hi, x_um = df_x    y_lm, y_hi, y_um = df_y    # 绘制顺序按大致的尺寸或视觉重点排列(UMICs底层, HICs中层, LMICs顶层)    ax.scatter(x_um, y_um, s=sizes[35:], c=colors_map['UMICs']['face'], alpha=0.55,               edgecolors=colors_map['UMICs']['edge'], linewidth=1.5, zorder=1)    ax.scatter(x_hi, y_hi, s=sizes[15:35], c=colors_map['HICs']['face'], alpha=0.55,               edgecolors=colors_map['HICs']['edge'], linewidth=1.5, zorder=2)    ax.scatter(x_lm, y_lm, s=sizes[:15], c=colors_map['LMICs']['face'], alpha=0.55,               edgecolors=colors_map['LMICs']['edge'], linewidth=1.5, zorder=3)    # 计算整体 r 值    all_x = np.concatenate([x_lm, x_hi, x_um])    all_y = np.concatenate([y_lm, y_hi, y_um])    r_val, _ = pearsonr(all_x, all_y)    # 添加 r 值 (斜体,位于右下角)    ax.text(0.96, 0.05, f"r = {r_val:.2f}", transform=ax.transAxes,            ha='right', va='bottom', fontsize=11, fontstyle='italic', zorder=5)    # 设置标题与坐标轴标签    ax.set_title(title_str, fontsize=11, pad=10)    ax.set_xlabel(x_label, fontsize=10)    if y_label:        ax.set_ylabel(y_label, fontsize=10)    # 边框与刻度美化:去除右侧、顶部边框,刻度线朝外 (direction='out')    ax.spines['top'].set_visible(False)    ax.spines['right'].set_visible(False)    ax.spines['left'].set_color('#333333')    ax.spines['bottom'].set_color('#333333')    ax.tick_params(axis='both', which='major', labelsize=9,                    colors='#333333', direction='out', length=4)# ================= 3. 多图表排版生成 =================# 创建 3行 x 3列的排版fig, axes = plt.subplots(3, 3, figsize=(16, 11))titles = [    "Density of chain outlets",    "Density of non-chain outlets",    "Ratio of non-chain to chain outlets",    "Percentage of grocery sales from\nchain outlets",    "Sales of unhealthy food per capita",    "Percentage of unhealthy food sales\nfrom chain outlets",    "Digital grocery sales per capita"]xlabels = [    "AAPC (%) for the density of chained outlets",    "AAPC (%) for the density of non-chain outlets",    "AAPC (%) for the ratio of non-chain to chain outlets",    "AAPC (%) of percentage of sales from chain outlets",    "AAPC (%) of unhealthy food sales (kg per capita)",    "AAPC (%) of percentage of unhealthy sales\nfrom chain outlets",    "AAPC (%) of digital grocery sales (US$ per capita)"]for i in range(9):    row, col = i // 3, i % 3    ax = axes[row, col]    if i < 7:  # 绘制前 7 个散点图        df_x, df_y, sizes = generate_group_data(i)        y_label = "AAPC of obesity prevalence (%)" if col == 0 else None        draw_bubble_plot(ax, df_x, df_y, sizes, titles[i], xlabels[i], y_label)    elif i == 7:  # 第 8 个格子绘制全局图例        ax.axis('off')        # 手动构建符合散点样式的完美图例句柄        legend_elements = [            Line2D([0], [0], marker='o', color='w', label='LMICs',                   markerfacecolor=colors_map['LMICs']['face'],                    markeredgecolor=colors_map['LMICs']['edge'],                   markersize=10, alpha=0.6, markeredgewidth=1.5),            Line2D([0], [0], marker='o', color='w', label='HICs',                   markerfacecolor=colors_map['HICs']['face'],                    markeredgecolor=colors_map['HICs']['edge'],                   markersize=10, alpha=0.6, markeredgewidth=1.5),            Line2D([0], [0], marker='o', color='w', label='UMICs',                   markerfacecolor=colors_map['UMICs']['face'],                    markeredgecolor=colors_map['UMICs']['edge'],                   markersize=10, alpha=0.6, markeredgewidth=1.5)        ]        # 放置无边框图例        ax.legend(handles=legend_elements, loc='center', frameon=False,                   fontsize=12, handletextpad=0.2, labelspacing=1.0)    else:          ax.axis('off')# 手动调整子图间距,为长标题和底部的坐标轴标签留出充分空间plt.subplots_adjust(hspace=0.65, wspace=0.25, left=0.05, right=0.98, top=0.92, bottom=0.08)plt.show()

这段代码为你解决了哪些排版痛点?

高级通透的色彩美学摒弃了死板的纯黑边框!代码中采用了自定义的字典映射,使用了“半透明面色 + 同色系深色边框”的组合。即使数据点密集堆叠,画面依然干净通透,层次分明(UMICs、HICs、LMICs 三组数据层层叠放,细节拉满)。

极致清爽的 3x3 空间布局巧妙利用 3x3 网格排版 7 张相关性图表。自动去除了每一行冗余的 Y 轴标签,只保留最左侧的标注,最大程度把空间留给数据本身。

专业独立的全局图例图例经常遮挡数据点?这段代码利用Line2D重构了图例句柄,并将其巧妙放置在第8个空白网格中。无边框设计配合精准的间距调整,让整体排版像杂志一样优雅。

满分的学术规范细节自动计算并添加了皮尔逊相关系数(r 值),并严格采用了斜体排版;刻度线统一设置为朝外(direction='out');同时去除了顶部和右侧的边框,完全符合国际顶刊的审美标准。

这种带权重的气泡散点图尤其适合公共卫生、环境流行病学、生态学等需要处理不同规模样本(如不同面积的国家、不同区域人口)的数据场景📦。

觉得有用的话,点个「在看」👍 或转发给同门,一起卷起来~

我们下一个小技巧见 👋😊

最新文章

随机文章