当前位置:首页>python>用纯 Python 手撸一个 AI Agent!!

用纯 Python 手撸一个 AI Agent!!

  • 2026-10-11 05:41:05
用纯 Python 手撸一个 AI Agent!!

大家好,我是小寒

今天我将带领大家用纯 Python 构建一个真正的智能体,无需任何智能体框架。

一旦你了解了智能体的本质,就会发现它出奇地简单。

它由一个while循环、一组工具、一个目标和一个停止条件组成。

这就是它的全部理念。

让我们构建一个真正有用的智能体——给定一个文本文件目录,它能找出哪个文件提到了某个主题并返回结果。

我们正在构建什么

我们的代理只有一个目标:“找到目录中与数据库索引相关的文件,并总结其内容。”

为了完成这个目标,它需要搜索目录、读取文件并进行推理。

因此,我们将为它提供三种工具:

  • list_files — 查看目录中的内容。
  • read_file — 读取文件内容。
  • finish — 宣布完成并交还答案。

最后一点比表面看起来更重要。

结束代理循环的一个简洁方法是给模型一个明确的 “我完成了” 工具,这样结束就成了一个刻意的行为,而不是措辞上的偶然。

思考→行动→观察循环

这是编写任何代码之前的思维模型。

每一步:

  1. 思考 —— 该模型会审视目标和迄今为止发生的一切,并决定下一步行动。
  2. 行动 — 如果它想要一个工具,它会发出一个结构化的工具调用(名称 + 参数)。
  3. 观察 —— 你的代码运行该工具,捕获结果,并将其反馈。

然后它开始循环。模型会再次思考,现在它知道了工具的返回结果。

它会一直循环,直到被调用finish——或者直到到达我们设置的预算。

预算就像一道安全屏障,防止混乱的代理无限循环并耗尽你的账户。

步骤 1:将工具定义为纯 Python 函数

工具只是函数。

智能体不会运行它们——只有当模型发出请求时,你才会运行它们。

保持工具简洁明了、可预测,并且始终返回模型可以读取的字符串。

import osNOTES_DIR = "notes"def list_files(directory: str) -> str:"""Return the file names in a directory, one per line."""    try:        names = sorted(os.listdir(directory))return"\n".join(names) if names else"(directory is empty)"    except OSError as e:return f"Error listing '{directory}': {e}"def read_file(path: str) -> str:"""Return the text contents of a file (truncated to keep tokens sane)."""    try:        with open(path, "r", encoding="utf-8") as f:            text = f.read()return text[:4000] if len(text) > 4000 else text    except OSError as e:return f"Error reading '{path}': {e}"# A registry so the loop can dispatch by name.TOOL_FUNCTIONS = {"list_files": list_files,"read_file": read_file,}

finish 它很特殊:它本身不做任何操作,只是发出完成信号。

我们将在循环中直接处理它,而不是在注册表中处理。

请注意,每个工具都会捕获自身的错误并以文本形式返回,当某个工具失败时,代理应该能够检测到该错误并进行相应的调整,而不是导致程序崩溃。

步骤 2:向模型描述工具

模型无法直接查看你的 Python 代码。

它需要 JSON Schema 描述,以便了解每个工具的名称、用途和参数。

编写描述时,要像向新同事介绍工作一样——说明何时使用该工具,而不仅仅是它的功能。

TOOLS = [    {"name": "list_files","description": "List the file names in a directory. Use this first to ""discover what files exist before reading any of them.","input_schema": {"type": "object","properties": {"directory": {"type": "string", "description": "Directory path to list"}            },"required": ["directory"],        },    },    {"name": "read_file","description": "Read the text contents of a single file. Use this to ""inspect a file you found with list_files.","input_schema": {"type": "object","properties": {"path": {"type": "string", "description": "Path to the file to read"}            },"required": ["path"],        },    },    {"name": "finish","description": "Call this when you have the answer. Provide the final ""answer to the user's goal in the 'answer' field.","input_schema": {"type": "object","properties": {"answer": {"type": "string", "description": "The final answer"}            },"required": ["answer"],        },    },]

步骤 3:代理循环本身

现在到了关键所在。

在这里,“思考→行动→观察” 形成了一个while有预算限制的循环。

import jsonfrom openai import OpenAIclient = OpenAI(    api_key="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",    base_url="https://openrouter.ai/api/v1")SYSTEM_PROMPT = ("You are a file-investigation agent. You achieve the user's goal by ""calling tools to explore the filesystem, then calling 'finish' with your ""answer. Explore before you conclude. Don't guess file contents — read them.")def run_agent(goal: str, step_budget: int = 8) -> str:# 1. 在 messages 开头加入 system prompt    messages = [        {"role": "system", "content": SYSTEM_PROMPT},        {"role": "user", "content": goal}    ]for step in range(1, step_budget + 1):print(f"\n--- step {step}/{step_budget} ---")# 2. 调用 API(移除了顶层 system 参数)        response = client.chat.completions.create(            model="openrouter/free",            max_tokens=2000,            tools=TOOLS,  # 注意:OpenAI 格式的 TOOLS 结构需要是 {"type": "function", "function": {...}}            messages=messages,        )        choice = response.choices[0]        message = choice.message# 将模型的回复追加到历史记录中(必须完整保存,包含 tool_calls)        messages.append(message)# 3. 如果模型没有发起工具调用,直接返回文本回答并结束if not message.tool_calls:return message.content or "(agent ended without a final answer)"# 4. 遍历处理每一个工具调用for tool_call in message.tool_calls:            func_name = tool_call.function.name# OpenAI 返回的参数是 JSON 字符串,需要解析为字典            args = json.loads(tool_call.function.arguments)# 特殊工具 finish:获取答案并退出if func_name == "finish":return args.get("answer", "")            func = TOOL_FUNCTIONS.get(func_name)if func is None:                result = f"Error: unknown tool '{func_name}'"else:print(f"  calling {func_name}({args})")                try:                    result = func(**args)                except Exception as e:                    result = f"Error executing '{func_name}': {e}"# 5. OpenAI 格式:每一个工具结果作为一条 role="tool" 的消息单独追加            messages.append({"role": "tool","tool_call_id": tool_call.id,"content": str(result)            })return"(step budget exhausted before the agent finished)"

逐步了解一次回合的生命周期

  • 我们调用模型时会传入目标、工具以及目前为止的所有对话记录。API 是无状态的——每次调用都会重新发送完整的历史记录。模型本身不会“记住”什么;是你提醒它。
  • 如果返回的 stop_reason 是 "end_turn",说明模型只是用普通文本做了回答,没用任何工具。我们直接返回这个文本并结束。
  • 如果模型决定用工具,千万注意!必须把模型返回的这段回复(里面带着 tool_use 结构)原封不动存进对话历史里。接着,你的代码去真正执行这些工具,并将结果封装成 tool_result 返还给模型。这里还有一个死规则:每一个 tool_use 指令,都必须有一个带有相同 tool_use_id 的 tool_result 结果与之对应,一个都不能少。
  • 一旦模型在调用工具时选择了 finish,说明它拿到最终答案了,直接跳出整个大循环,把答案吐给用户。
  • 如果跑着跑着达到了我们设定的最大步数,程序就会优雅地强制退出,防止 Agent 在里面打转陷入死循环。

第四步:一个可验证的目标

一个好的智能体任务应该有一个可验证的目标——一个你可以验证其真实性的目标,而不仅仅是听起来合情合理。

我们的目标就是可验证的:我们可以打开智能体指定的文件,确认它确实在讨论索引问题。

让我们创建测试数据并运行它。

if __name__ == "__main__":    os.makedirs(NOTES_DIR, exist_ok=True)    with open(f"{NOTES_DIR}/standup.txt", "w") as f:        f.write("Discussed the Q3 roadmap and the hiring freeze. No blockers.")    with open(f"{NOTES_DIR}/perf.txt", "w") as f:        f.write("Slow queries traced to a missing B-tree index on orders.user_id. ""Adding the index cut p95 latency from 800ms to 40ms.")    with open(f"{NOTES_DIR}/lunch.txt", "w") as f:        f.write("Team voted for tacos. Again.")    answer = run_agent(        f"Find which file in the '{NOTES_DIR}' directory talks about database "        f"indexing, and summarize what it says.",        step_budget=8,    )print("\n=== FINAL ANSWER ===")print(answer)

运行它,你会看到跟踪信息:代理程序调用了  list_files("notes") ,发现了三个文件,读取了全部文件,然后返回 finish 类似这样的信息。

perf.txt 讨论了数据库索引——缺少 B 树索引,导致查询速度缓慢;添加索引后,p95 延迟从 800 毫秒降至 40 毫秒。

最后

—

今天的分享就到这里。如果觉得近期的文章不错,请点赞,转发安排起来。‍‍
欢迎大家进高质量 python 学习群‍‍‍‍‍‍‍‍

「进群方式:加我微信,备注 “python”」

往期回顾

Fashion-MNIST 服装图片分类-Pytorch实现

python 探索性数据分析(EDA)案例分享

深度学习案例分享 | 房价预测 - PyTorch 实现

万字长文 |  面试高频算法题之动态规划系列

面试高频算法题之回溯算法(全文六千字)  

如果对本文有疑问可以加作者微信直接交流。

最新文章

随机文章