当前位置:首页>python>【CMIP6数据下载】 wget或python

【CMIP6数据下载】 wget或python

  • 2026-10-11 08:22:38
【CMIP6数据下载】 wget或python

最近在下载 CMIP6 数据时,遇到了 ESGF wget 脚本无法正常下载的问题。原始数据是 CNRM-CM6-1 模式的 historical 情景下 tauu 变量,下载脚本来自 ESGF 平台。整个过程从认证失败、节点连接失败、校验失败,最终改用 Python 在 Jupyter 中稳定完成下载。

这篇文章记录完整排查过程和最终可运行代码。

数据下载网址

https://wcrp-cmip.org/cmip-data-access/

此网站给出CMIP6数据有关的说明文件以及下载网址

点击here,进入数据下载页面

https://esgf-node.ornl.gov/search

此页面由于网络不稳定,可能会打不开,多次尝试即可

  1. 进入页面之后可在左侧对数据进行筛选,回车之后右侧列表会进行刷新,点击需要的数据后面的下载按钮,会下载包含wget命令的.sh文件

  2. 以CNRM-CM6-1-HR.historical.r1i1p1f2.Amon.tauu.gr.sh 为例

  3. 在linux系统里cd到此sh文件所在文件夹下输入

    ./CNRM-CM6-1-HR.historical.r1i1p1f2.Amon.tauu.gr.sh

    回车,即开始下载

最初的问题:wget 脚本无法下载

运行

./CNRM-CM6-1.historical.r1i1p1f2.Amon.tauu.gr.sh -H

脚本可以启动,但是提示:

No ESG Credentials foundJava could not be foundmyproxy-logon could not be found

这说明问题主要出在 ESGF 认证环境上。旧版 wget 脚本会尝试使用证书认证,而本地并没有配置 ESGF 证书、Java 和 myproxy 工具。

跳过认证:使用 -s -i

./CNRM-CM6-1.historical.r1i1p1f2.Amon.tauu.gr.sh -s -i

其中-s  跳过认证-i  不检查服务器证书

这一步成功绕过了 ESGF 认证流程,说明认证不再是主要障碍。

新问题:CEDA 数据节点连接失败

随后脚本开始连接数据节点:

https://esgf.ceda.ac.uk/thredds/fileServer/...

但出现:

failed: Resource temporarily unavailable

这说明问题已经不是认证,而是当前网络连接 CEDA 节点不稳定。CMIP6 下载经常会遇到这种情况:ESGF 搜索节点可以打开,但真正存放数据的数据节点不一定能稳定连接。

改用 Python 下载

为了绕开 wget 脚本中的认证和证书流程,可以直接用 Python 解析 .sh 文件中的下载链接,然后用 requests 下载数据。

经过调试,下面是最终成功运行的脚本,只需要修改前面的路径即可运行

import reimport hashlibimport requestsfrom pathlib import Pathfrom time import sleepimport os# ===== 修改为自己的工作目录 =====work_dir = Path("/path/to/your/cmip6_download_folder")os.chdir(work_dir)print(os.getcwd())# ===== 修改为自己的 ESGF wget 脚本路径 =====sh_file = work_dir / "CNRM-CM6-1.historical.r1i1p1f2.Amon.tauu.gr.sh"# ===== 数据保存目录 =====save_dir = work_dir# ===============================requests.packages.urllib3.disable_warnings()chunk_size = 1024 * 1024max_retry = 10# 读取 ESGF wget 脚本text = sh_file.read_text(encoding="utf-8", errors="ignore")# 提取文件名、下载链接、校验类型和校验值pattern = r"'([^']+\.nc)'\s+'(https://[^']+)'\s+'([^']+)'\s+'([^']+)'"files = re.findall(pattern, text)print(f"Found {len(files)} file(s)")def sha256sum(filename):    """计算本地文件的 SHA256 值"""    h = hashlib.sha256()    with open(filename, "rb") as f:        for block in iter(lambda: f.read(1024 * 1024), b""):            h.update(block)    return h.hexdigest()def is_netcdf_file(filename):    """检查文件是否为 NetCDF 格式"""    with open(filename, "rb") as f:        head = f.read(8)    return head.startswith(b"CDF") or head.startswith(b"\x89HDF")for fname, url, chk_type, chk_value in files:    out = save_dir / fname    tmp = save_dir / (fname + ".tmp")    print("\n======================================")    print(fname)    print(url)    # 删除旧文件,避免上一次错误下载影响本次结果    if out.exists():        print("Removing old file")        out.unlink()    if tmp.exists():        tmp.unlink()    for attempt in range(1, max_retry + 1):        try:            print(f"\nAttempt {attempt}")            with requests.get(                url,                stream=True,                timeout=(30, 300),                verify=False,                allow_redirects=True            ) as r:                print("HTTP status:", r.status_code)                print("Content-Length:", r.headers.get("Content-Length"))                if r.status_code != 200:                    raise RuntimeError(f"HTTP status {r.status_code}")                total = 0                with open(tmp, "wb") as f:                    for chunk in r.iter_content(chunk_size=chunk_size):                        if chunk:                            f.write(chunk)                            total += len(chunk)                            # 每下载约 10 MB 输出一次进度                            if total % (10 * 1024 * 1024) < chunk_size:                                print(f"Downloaded: {total/1024**2:.1f} MB")                print(f"Downloaded size: {total/1024**2:.2f} MB")            # 检查是否为 NetCDF 文件            if not is_netcdf_file(tmp):                print("Not NetCDF file, removing")                tmp.unlink()                raise RuntimeError("Invalid file")            # 下载完成后改为正式文件名            tmp.rename(out)            # SHA256 校验            if chk_type.lower() == "sha256":                print("Checking SHA256...")                local_hash = sha256sum(out)                if local_hash.lower() == chk_value.lower():                    print("SHA256 OK")                    break                else:                    print("SHA256 FAILED")                    print("local :", local_hash)                    print("remote:", chk_value)                    out.unlink()                    raise RuntimeError("Checksum mismatch")            break        except Exception as e:            print("Error:", e)            sleep(15)    else:        print("FAILED after all retries")

成功标志

如果代码正常运行,最后会看到:

Download finishedChecking SHA256...SHA256 OK

这说明:

  1. 文件完整下载;

  2. 文件格式是 NetCDF;

  3. 本地文件和 ESGF 提供的 SHA256 校验值一致;

  4. 数据可以安全用于后续分析。

最新文章

随机文章