ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

OpenGlass 多模型图像描述流水线解析:以 prompts/series_1/img_23.md 为例

OpenGlass 多模型图像描述流水线解析:以 prompts/series_1/img_23.md 为例 人工智能AI 应用智能硬件本地部署可穿戴AI Agent【免费下载链接】OpenGlassTurn any glasses into AI-powered smart glasses项目地址https://gitcode.com/GitHub_Trending/op/OpenGlass点击查看免费下载导读本文以 OpenGlass 仓库中的 prompts/series_1/img_23.md 为样本完整拆解该项目一张图片 → 四个视觉语言模型并行描述 → 结构化 Markdown 报告的生成流水线。读完本文你将掌握 prompts/generate.ts 的批处理机制、sources/agent/imageDescription.ts 的提示词设计以及 moondream、llava 系列模型在同一输入下的输出差异并了解这些描述如何支撑智能眼镜端的多模态理解能力。一、img_23.md 是什么一份四模型对比描述报告img_23.md是 OpenGlass 仓库 prompts/series_1 目录下 57 份同类样本之一其本质是同一张图片在不同视觉语言模型VLM下的描述输出记录。该文件由四段构成每段以####标题####为分隔标记####Description#### 默认模型描述 ####Description (llava-llama3)#### llava-llama3 描述 ####Description (llava:34b-v1.6)#### llava:34b-v1.6 描述 ####Description (moondream:1.8b-v2-fp16)#### moondream 描述与它配套的图片是 prompts/series_1/img_23.jpeg600x800 尺寸。整份报告不需要人工撰写——它是由 prompts/generate.ts 自动批量生成的产物这一点从文件命名img_N.md与img_N.jpeg一一对应和 package.json 中的prompts: ts-node ./prompts/generate.ts脚本即可印证。从仓库结构看prompts/目录扮演着数据采集与评测集的角色它既为开发团队沉淀了一批真实场景的图像 多模型描述样本也为后续基于描述文本的问答 Agent如llamaFind提供了可复用的语料。二、生成流水线prompts/generate.ts 的批处理机制prompts/generate.ts 是整个描述报告的生产入口核心流程分为三步2.1 扫描并收集图片样本脚本读取prompts/下所有子目录series_1、series_2等遍历其中所有.jpeg文件并把输出路径规整为同名的.md文件let allFiles fs.readdirSync(__dirname); for (let f of allFiles) { if (fs.statSync(path.join(__dirname, f)).isDirectory()) { let files fs.readdirSync(path.join(__dirname, f)); for (let s of files) { if (s.endsWith(.jpeg)) { let image fs.readFileSync(path.join(__dirname, f, s)); imageTests.push({ path: path.join(__dirname, f, s).replace(.jpeg, .md), image, outputs: }); } } } }可以看到img_23.md与img_23.jpeg正是通过replace(.jpeg, .md)这一规则绑定的。2.2 依次跑四组描述测试runTest是一个通用执行器为每个样本调用一次传入的测试函数并把返回结果以####标题####前缀追加到该样本的outputs中同时用cli-progress渲染进度条。脚本依次注册了四个测试await runTest(Description, async (img) { // 使用默认模型 moondream:1.8b-v2-fp16 return await imageDescription(img); }); await runTest(Description (llava-llama3), async (img) { return await imageDescription(img, llava-llama3); }); await runTest(Description (llava:34b-v1.6), async (img) { return await imageDescription(img, llava:34b-v1.6); }); await runTest(Description (moondream:1.8b-v2-fp16), async (img) { return await imageDescription(img, moondream:1.8b-v2-fp16); });这解释了img_23.md中四段的由来第一个Description段实际用的是imageDescription的默认参数moondream:1.8b-v2-fp16因此它与最后一段调用了同一个模型而中间两段分别来自 llava-llama3 与 llava:34b-v1.6。这也提醒读者报告中默认段与moondream 段内容可能高度相似属于预期行为。值得注意的还有被注释掉的模糊检测测试// console.log(Run blurry tests); // for (let i of imageTests) { // i.outputs ####Blurry####\n; // i.outputs await imageBlurry(i.image) \n; // }对应 sources/agent/imageBlurry.ts 中的imageBlurry函数——它可用于在批量描述前过滤失焦照片目前处于实验阶段未写入最终报告。2.3 写回文件全部测试完成后脚本将累计的outputs写回对应的.md路径。因此img_23.md的每一行内容都可视为一次真实模型调用的原始返回是研究 VLM 行为差异的一手数据。运行方式仓库 package.json# 使用 npm npm run prompts # 或使用 yarn yarn prompts三、核心 AgentimageDescription 的提示词设计描述请求由 sources/agent/imageDescription.ts 中的imageDescription函数发出export async function imageDescription(src: Uint8Array, model: KnownModel moondream:1.8b-v2-fp16): Promisestring { return ollamaInference({ model: model, messages: [{ role: system, content: You are a very advanced model and your task is to describe the image as precisely as possible. Transcribe any text you see. }, { role: user, content: Describe the scene, images: [src], }] }); }该函数揭示了三个关键设计点System 提示词强调精确描述 转录可见文字describe the image as precisely as possible和Transcribe any text you see直接决定了模型输出风格——偏客观、偏细节这正是后续问答 Agent 需要高质量事实来源的原因。User 提示词极简仅Describe the scene把自由发挥空间留给模型便于对比不同模型的语言组织能力。图片以images: [src]数组内联传递走的是 Ollama 原生多模态消息格式而非 base64 文本拼接。同文件中还定义了基于描述文本的问答函数llamaFind走 sources/modules/groq-llama3.ts 的groqRequest模型为llama3-70b-8192和openAIFind走 sources/modules/openai.ts 的gptRequest模型为gpt-4o。两者的 system 提示词一致地要求只依据提供的描述回答问题、不得泛化、不得提及图片本身、回答要简洁具体。也就是说img_23.md这类描述报告正是这些问答 Agent 的证据库。四、底层实现Ollama 推理、退避重试与文本规整imageDescription最终调用 sources/modules/ollama.ts 的ollamaInferenceexport type KnownModel | llama3 | llama3-gradient | llama3:8b-instruct-fp16 | llava-llama3 | llava:34b-v1.6 | moondream:1.8b-v2-fp16 | moondream:1.8b-v2-moondream2-text-model-f16KnownModel联合类型限定了可选模型清单imageDescription的model参数必须属于该集合。推理请求的核心逻辑const response await backoffany(async () { let converted: { role: string, content: string, images?: string[] }[] []; for (let message of args.messages) { converted.push({ role: message.role, content: trimIdent(message.content), images: message.images ? message.images.map((image) toBase64(image)) : undefined, }); } let resp await axios.post(keys.ollama, { stream: false, model: args.model, messages: converted, }); return resp.data; }); return trimIdent(((response.message?.content ?? ) as string));4.1 请求格式stream: false关闭流式输出等待完整响应后再返回messages[].images将Uint8Array经 sources/utils/base64.ts 的toBase64转为 base64 字符串imageDescription直接传原始字节不携带data:image/jpeg;base64,前缀而 sources/app/DeviceView.tsx 中的toBase64Image才会拼接该 MIME 前缀用于前端展示keys.ollama来自 sources/keys.ts即环境变量EXPO_PUBLIC_OLLAMA_API_URLREADME 给出的典型值为http://localhost:11434/api/chat。4.2 退避重试backoff来自 sources/utils/time.ts 的createBackoff默认参数为minDelay250ms、maxDelay1000ms、maxFailureCount50 次失败时按随机指数退避策略等待后重试。这意味着批量跑 57 张图时单次 Ollama 服务波动不会导致整个流水线中断。4.3 文本规整sources/utils/trimIdent.ts 的trimIdent负责两件事去掉首尾空行、并移除所有非空行的公共缩进。由于 system/user 提示词在源码中使用了多行模板字符串trimIdent保证了发送给模型的文本与最终落盘的响应文本都干净规整——这也是img_23.md中描述段落没有多余缩进的原因。4.4 模型准备按 README.md 的说明使用前需先在本地拉取默认视觉模型ollama pull moondream:1.8b-v2-fp16llava 系模型llava-llama3、llava:34b-v1.6同样通过ollama pull准备即可并在调用时作为model参数传入。五、案例深读四份描述眼中的 img_23现在回到文档本体。以下四段描述均出自 img_23.md描述的是同一张 600x800 的室内建筑照片但各自的观察重点差异鲜明。5.1 默认模型moondream:1.8b-v2-fp16——整体空间The image features a large, open space with an industrial-style ceiling. The ceiling is made of metal and has a grid pattern on it, giving the area a unique appearance. A person can be seen in the background, possibly observing or working within this space.侧重宏观场景判定为工业风格天花板 金属网格图案的大型开放空间并注意到背景中有一位可能在观察或工作的人。信息密度适中是典型的场景级概括。5.2 llava-llama3——色彩、几何与人物The image captures a unique perspective of the interior of an architectural structure. The dominant color in the frame is yellow, which is the exterior wall of the building. This wall is made up of small square panels arranged in a grid-like pattern, giving it a sense of uniformity and order.Two large metal beams run diagonally across the image from the top left to the bottom right corner. These beams, coupled with the triangular shape of the roof, create an interesting geometric contrast against the yellow wall.On the right side of the image, theres a blurred face partially visible in the background. This adds a touch of mystery and intrigue to the scene, making one wonder about the identity of this person and what they might be doing in this setting.信息量最大捕捉到主色调黄色外墙、方形小面板的网格纹理、两根自左上向右下对角贯穿的金属大梁、三角形屋顶以及右侧背景中模糊的人脸。与前一段相比llava-llama3 明显更关注构图几何关系与叙事性细节甚至对模糊人脸展开了推测性描写。5.3 llava:34b-v1.6——机位与光影The image appears to be taken from a low angle, looking upwards towards what could be the interior of a building with an exposed metal framework. The framework consists of rafters or beams that intersect at various angles, forming a geometric pattern against the sky. There are no texts visible in this image. The lighting suggests its daytime as sunlight is streaming through the structure, casting shadows and highlighting the metallic surfaces. It seems like a candid shot capturing an architectural detail rather than a posed composition.关注点转向拍摄视角与光影判断为低角度仰拍、暴露的金属框架与桁架交错、阳光穿过结构投射阴影并明确给出图中无文字白天光线更像抓拍的建筑细节而非摆拍等判断。34B 参数的模型展现了更强的场景推理与元描述能力。5.4 moondream:1.8b-v2-fp16显式指定——尺度感The image features a high ceiling with metal beams that are crossed, creating an interesting architectural design. A person can be seen looking down from the top of the building through these beams, adding a sense of scale and perspective to the scene. The background is not described in detail but adds depth to the overall composition, suggesting the presence of other buildings or structures nearby.与前一段同为 moondream 模型但对人物的描述更进一步不仅看到人还推断其正透过钢梁向下看并因此强化了画面的尺度与透视感。这也验证了第二节的判断——默认段与显式指定段使用同一模型输出风格一致仅在措辞上略有差异。5.5 对比小结模型侧重维度亮点信息语言风格moondream:1.8b-v2-fp16默认整体空间工业金属网格天花板、背景人物概括、中性llava-llama3色彩与几何黄色主调、对角钢梁、三角形屋顶、模糊人脸细节丰富、带叙事性llava:34b-v1.6机位与光影低角度仰拍、白天光照、无文字判定推理型、元描述moondream:1.8b-v2-fp16显式尺度与透视人物向下看、画面尺度感概括、中性这份对比本身就是多模型评测的价值所在不同参数量与训练数据的模型对同一场景的注意焦点天然不同。在 OpenGlass 的语境下这意味着多模型并行描述比单一模型更能覆盖结构细节 人物活动 光影环境等不同信息维度为上层问答 Agent 提供更完整的证据集。六、从描述到智能眼镜描述文本的上游与下游6.1 上游照片从眼镜到云端img_23.jpeg这类图片在真实产品链路中来自 sources/app/DeviceView.tsx 的usePhotos通过 Web Bluetooth 订阅特征值19b10005-e8f2-537e-4f6c-d104768a1214按数据包 ID 拼装 JPEG 分片遇到0xff 0xff边界标记即认为一张照片接收完毕随后经rotateImage旋转 270 度后进入照片队列同时向控制特征值19b10006-...写入0x05启动每 5 秒自动拍照。6.2 下游描述驱动的问答与语音照片队列会通过InvalidateSync批量交给Agentagent.addPhoto模型描述结果经agentState.answer呈现并在产生答案时调用 sources/modules/openai.ts 的textToSpeech播报tts-1模型、nova音色。而Agent内部正是依赖llamaFind/openAIFind这类读描述、答问题的函数——所以img_23.md这样的描述报告既是评测数据也是问答系统的推理素材。七、从源码结构可推断的扩展方向结合仓库现状可以推断该描述流水线尚有以下扩展空间均以当前仓库代码为准非已实现能力模糊过滤sources/agent/imageBlurry.ts 的imageBlurry已存在但未启用可在批量描述前剔除失焦照片降低无效推理开销跨模型投票/融合既然同一图片已有四路描述未来可基于多路输出做一致性融合或分歧检测用于识别模型幻觉描述语料反哺评测prompts/series_1下的 57 组样本天然构成一个图文描述基准集可用于回归对比新版本 VLM 的稳定性。八、小结img_23.md并非孤立文件而是 OpenGlass图像采集 → 多模型并行描述 → 结构化落盘 → 问答/语音输出整条多模态链路的缩影。通过 prompts/generate.ts 的批处理、sources/agent/imageDescription.ts 的提示词设计、sources/modules/ollama.ts 的推理封装以及KnownModel模型矩阵开发者可以用一条命令复现整套描述报告并以四路输出对比为抓手持续评估 moondream 与 llava 系列模型在真实眼镜场景下的表现力差异。赞分享人工智能AI 应用智能硬件本地部署可穿戴AI Agent【免费下载链接】OpenGlassTurn any glasses into AI-powered smart glasses项目地址https://gitcode.com/GitHub_Trending/op/OpenGlass点击查看免费下载相关推荐OpenGlass 多模态图片描述流水线解析以 prompts/series_1/img_2.md 为例OpenGlass 多模态图片描述流水线解析以 prompts/series_1/img_2.md 为例 导读 本文以 OpenGlass 仓库中 promp人工智能AI 应用智能硬件本地部署可穿戴AI AgentOpenGlass 图像描述评测以 prompts/series_1/img_15.md 为例解析多模型视觉输出对比OpenGlass 图像描述评测以 prompts/series_1/img_15.md 为例解析多模型视觉输出对比 导读 prompts/series_1/人工智能AI 应用智能硬件本地部署可穿戴AI AgentOpenGlass 视觉模型图像描述评测解析以 prompts/series_1/img_12.md 为例的多模态模型对比OpenGlass 视觉模型图像描述评测解析以 prompts/series_1/img_12.md 为例的多模态模型对比 OpenGlass 是一款将普通眼人工智能AI 应用智能硬件本地部署可穿戴AI Agent上一篇Simple Keyboard一款轻量级、高度可定制的虚拟键盘库下一篇RuoYi-AI企业级智能平台高性能分布式AI应用开发完整方案创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表