ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

DeepSeek-Agent-Harness-2026终极指南-第9章第42节-核心工具集开发-globgrep:给Agent装上代码检索的眼睛

DeepSeek-Agent-Harness-2026终极指南-第9章第42节-核心工具集开发-globgrep:给Agent装上代码检索的眼睛 DeepSeek Agent Harness 2026终极指南 - 第9章第42节 glob/grep给Agent装上代码检索的眼睛第41节的bash执行器让Agent能跑命令了但还差一个能力——找代码。你让它找一下所有用到get_weather的地方它只能一个个文件读效率很低。这节做代码检索工具glob和grepglob按文件名匹配grep按内容搜索。从此Agent能秒级定位代码不用一个个文件翻。本文导航为什么Agent需要代码检索glob按文件名匹配grep按内容搜索结果数量限制防止上下文爆炸目录排除跳过node_modules和.git完整实现search_tools.py实测Agent自主找代码小结为什么Agent需要代码检索想象一个场景你让Agent把所有用到get_weather的地方改成get_city_weather。没有检索工具Agent只能读tools.py找到get_weather读agent_loop.py看有没有用到读main.py看有没有用到…一个个文件翻如果有100个文件Agent要读100次效率极低。有了检索工具Agent可以grep get_weather秒级找到所有用到的地方逐个edit_file修改代码检索是Agent的眼睛让它能快速定位代码而不是盲人摸象。glob按文件名匹配glob是Unix命令按文件名模式匹配文件。Python的pathlib.Path.glob()实现了同样功能。典型用法*.py所有Python文件**/*.md所有Markdown文件递归test_*.py所有测试文件实现frompathlibimportPathtooldefglob(pattern:str,cwd:str.)-str: 按文件名模式匹配文件。 参数 - pattern: glob模式如 *.py、**/*.md - cwd: 工作目录默认当前目录 返回匹配的文件列表每行一个 cwd_pathPath(cwd).resolve()try:matcheslist(cwd_path.glob(pattern))ifnotmatches:returnf未找到匹配 {pattern} 的文件# 转为相对路径方便阅读result_lines[]forminmatches[:100]:# 最多100个rel_pathm.relative_to(cwd_path)result_lines.append(str(rel_path))result\n.join(result_lines)iflen(matches)100:resultf\n\n[共{len(matches)}个匹配已显示前100个]returnresultexceptExceptionase:returnf错误glob失败 -{str(e)}关键点Path.glob(pattern)返回匹配的文件列表**递归匹配**/*.py匹配所有子目录下的.py文件相对路径m.relative_to(cwd_path)转为相对路径方便阅读最多100个防止匹配太多占满上下文grep按内容搜索grep是Unix命令按正则表达式搜索文件内容。Python没有内置grep但可以用re模块实现。典型用法grep def test_搜索所有测试函数grep import.*requests搜索导入requests的地方grep -i error忽略大小写搜索error实现importretooldefgrep(pattern:str,path:str.,ignore_case:boolFalse,max_results:int50)-str: 按正则表达式搜索文件内容。 参数 - pattern: 正则表达式模式 - path: 搜索路径文件或目录默认当前目录 - ignore_case: 是否忽略大小写默认False - max_results: 最大结果数默认50 返回匹配结果格式为文件:行号: 内容 search_pathPath(path).resolve()# 编译正则flagsre.IGNORECASEifignore_caseelse0try:regexre.compile(pattern,flags)exceptre.errorase:returnf错误无效的正则表达式 -{str(e)}results[]# 如果是文件直接搜索ifsearch_path.is_file():files_to_search[search_path]else:# 如果是目录递归搜索所有文本文件files_to_search[fforfinsearch_path.rglob(*)iff.is_file()]forfile_pathinfiles_to_search:# 跳过二进制文件和隐藏目录ifany(part.startswith(.)forpartinfile_path.parts):continueiffile_path.suffixin{.pyc,.pyo,.so,.dll,.exe,.jpg,.png,.gif}:continuetry:withopen(file_path,r,encodingutf-8,errorsignore)asf:forline_num,lineinenumerate(f,start1):ifregex.search(line):rel_pathfile_path.relative_to(search_path.parent)ifsearch_path.is_dir()elsefile_path.name results.append(f{rel_path}:{line_num}:{line.rstrip()})iflen(results)max_results:breakiflen(results)max_results:breakexceptException:# 跳过无法读取的文件continueifnotresults:returnf未找到匹配 {pattern} 的内容result\n.join(results)iflen(results)max_results:resultf\n\n[已达到最大结果数{max_results}可能还有更多匹配]returnresult关键点re.compile(pattern)编译正则支持复杂模式re.IGNORECASE忽略大小写递归搜索search_path.rglob(*)递归遍历目录跳过二进制文件检查文件扩展名跳过隐藏目录.git、.venv等errorsignore跳过编码错误的文件格式文件:行号: 内容方便定位结果数量限制防止上下文爆炸glob和grep可能返回大量结果glob(**/*.py大项目可能有几千个Python文件grep import可能匹配几千行如果不限制结果会占满上下文窗口。解决方案glob最多100个结果grep默认50个结果可配置max_results超出部分截断并提示还有更多匹配。目录排除跳过node_modules和.git大项目有很多不需要搜索的目录.gitGit仓库node_modulesNode.js依赖.venvPython虚拟环境__pycache__Python缓存dist、build构建输出解决方案在遍历时跳过这些目录。# 跳过的目录名SKIP_DIRS{.git,node_modules,.venv,__pycache__,dist,build,.pytest_cache}forfile_pathinsearch_path.rglob(*):# 检查是否在跳过目录内ifany(partinSKIP_DIRSforpartinfile_path.parts):continue...完整实现search_tools.py把以上逻辑整合成完整模块# deep_pilot/search_tools.py —— 代码检索工具 v0.4from__future__importannotationsimportrefrompathlibimportPathfromdeep_pilot.tool_registryimporttool# 跳过的目录SKIP_DIRS{.git,node_modules,.venv,__pycache__,dist,build,.pytest_cache,.tox}# 跳过的文件扩展名SKIP_EXTS{.pyc,.pyo,.so,.dll,.exe,.jpg,.jpeg,.png,.gif,.bmp,.pdf,.zip,.tar,.gz}tooldefglob(pattern:str,cwd:str.)-str: 按文件名模式匹配文件。 参数 - pattern: glob模式如 *.py、**/*.md、test_*.py - cwd: 工作目录默认当前目录 返回匹配的文件列表每行一个最多100个 cwd_pathPath(cwd).resolve()try:matches[]formincwd_path.glob(pattern):# 跳过隐藏目录和跳过目录ifany(part.startswith(.)orpartinSKIP_DIRSforpartinm.parts):continuematches.append(m)iflen(matches)100:breakifnotmatches:returnf未找到匹配 {pattern} 的文件# 转为相对路径result_lines[]forminmatches:rel_pathm.relative_to(cwd_path)result_lines.append(str(rel_path))result\n.join(result_lines)iflen(matches)100:resultf\n\n[已达到最大结果数 100可能还有更多匹配]returnresultexceptExceptionase:returnf错误glob失败 -{str(e)}tooldefgrep(pattern:str,path:str.,ignore_case:boolFalse,max_results:int50)-str: 按正则表达式搜索文件内容。 参数 - pattern: 正则表达式模式如 def test_、import.*requests - path: 搜索路径文件或目录默认当前目录 - ignore_case: 是否忽略大小写默认False - max_results: 最大结果数默认50 返回匹配结果格式为文件:行号: 内容 search_pathPath(path).resolve()# 编译正则flagsre.IGNORECASEifignore_caseelse0try:regexre.compile(pattern,flags)exceptre.errorase:returnf错误无效的正则表达式 -{str(e)}results[]# 确定要搜索的文件列表ifsearch_path.is_file():files_to_search[search_path]else:files_to_search[]forfinsearch_path.rglob(*):ifnotf.is_file():continue# 跳过隐藏目录和跳过目录ifany(part.startswith(.)orpartinSKIP_DIRSforpartinf.parts):continue# 跳过二进制文件iff.suffix.lower()inSKIP_EXTS:continuefiles_to_search.append(f)# 搜索每个文件forfile_pathinfiles_to_search:try:withopen(file_path,r,encodingutf-8,errorsignore)asf:forline_num,lineinenumerate(f,start1):ifregex.search(line):# 计算相对路径ifsearch_path.is_dir():rel_pathfile_path.relative_to(search_path)else:rel_pathfile_path.name results.append(f{rel_path}:{line_num}:{line.rstrip()})iflen(results)max_results:breakiflen(results)max_results:breakexceptException:# 跳过无法读取的文件continueifnotresults:returnf未找到匹配 {pattern} 的内容result\n.join(results)iflen(results)max_results:resultf\n\n[已达到最大结果数{max_results}可能还有更多匹配]returnresult实测Agent自主找代码在deep_pilot/tools.py里导入检索工具# deep_pilot/tools.py —— v0.4 加入检索工具fromdeep_pilot.file_toolsimportread_file,write_file,edit_filefromdeep_pilot.bash_toolsimportrun_bashfromdeep_pilot.search_toolsimportglob,grep# 保留之前的工具tooldefget_weather(city:str)-str:获取指定城市今天的天气信息。returnf{city}晴天28°C实测Agent自主找代码uv run python-c from deep_pilot.agent_loop import run # 测试1找所有Python文件 print( 测试1找所有Python文件 ) answer run(找一下项目里所有的 .py 文件) print(f\nAgent回答: {answer}) print() # 测试2搜索特定函数 print( 测试2搜索 get_weather 函数 ) answer run(找一下代码里所有用到 get_weather 的地方) print(f\nAgent回答: {answer}) print() # 测试3搜索测试函数 print( 测试3搜索所有测试函数 ) answer run(找一下所有的测试函数def test_) print(f\nAgent回答: {answer}) 控制台输出精简 测试1找所有Python文件 2026-09-12 23:00:01 | INFO | agent_loop | Loop 第 1 轮 ↻ 2026-09-12 23:00:01 | INFO | agent_loop | → 调用工具: glob({pattern: **/*.py}) 2026-09-12 23:00:01 | INFO | agent_loop | ← 工具结果: deep_pilot/__init__.py deep_pilot/agent_loop.py deep_pilot/bash_tools.py deep_pilot/client.py deep_pilot/file_tools.py deep_pilot/search_tools.py deep_pilot/tools.py test_math.py 2026-09-12 23:00:01 | INFO | agent_loop | Loop 第 2 轮 ↻ Agent回答: 项目里有8个Python文件 - deep_pilot/ 目录下7个模块 - test_math.py 测试文件 测试2搜索 get_weather 函数 2026-09-12 23:00:02 | INFO | agent_loop | Loop 第 1 轮 ↻ 2026-09-12 23:00:02 | INFO | agent_loop | → 调用工具: grep({pattern: get_weather}) 2026-09-12 23:00:02 | INFO | agent_loop | ← 工具结果: deep_pilot/tools.py:12: def get_weather(city: str) - str: deep_pilot/tools.py:14: 获取指定城市今天的天气信息。 2026-09-12 23:00:02 | INFO | agent_loop | Loop 第 2 轮 ↻ Agent回答: get_weather 函数定义在 deep_pilot/tools.py 的第12-14行。 测试3搜索所有测试函数 2026-09-12 23:00:03 | INFO | agent_loop | Loop 第 1 轮 ↻ 2026-09-12 23:00:03 | INFO | agent_loop | → 调用工具: grep({pattern: def test_}) 2026-09-12 23:00:03 | INFO | agent_loop | ← 工具结果: test_math.py:5: def test_add(): 2026-09-12 23:00:03 | INFO | agent_loop | Loop 第 2 轮 ↻ Agent回答: 找到1个测试函数test_math.py 第5行的 test_add()。三个测试都通过了找Python文件Agent调glob(**/*.py)秒级找到8个文件搜索函数Agent调grep(get_weather)找到定义位置搜索测试函数Agent调grep(def test_)找到测试函数注意Agent的回答都是结构化的它理解了检索结果然后用清晰的方式呈现。小结代码检索是Agent的眼睛glob按文件名匹配grep按内容搜索秒级定位代码。Path.glob(pattern)支持*、**等通配符递归匹配文件。re.compile(pattern)支持正则表达式强大灵活。结果数量限制glob最多100个grep默认50个防止上下文爆炸。目录排除跳过.git、node_modules、.venv等不需要搜索的目录。二进制文件跳过检查扩展名跳过图片、编译文件等。格式友好grep返回文件:行号: 内容方便定位。DeepPilot v0.4检索工具完成——Agent能快速找代码了从盲人摸象升级到火眼金睛。下节预告文件、bash、检索三件套都齐了Agent已经能读改写代码、跑命令、找代码。但还差一个能力——联网。你让它查一下DeepSeek最新文档它只能说我没有联网能力。下一节做联网工具web_search和web_fetch搜索网页、抓取内容。从此Agent能查文档、查API、查最新信息真正成为你的全能助手。如果觉得本文对你有帮助欢迎点赞、收藏、关注三连本系列持续更新中关注不迷路~
返回列表