ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

Zeek SumStats 框架非集群实现解析:基于 non-cluster.zeek 的 epoch 处理与阈值检测机制

Zeek SumStats 框架非集群实现解析:基于 non-cluster.zeek 的 epoch 处理与阈值检测机制 网络安全网络IDS【免费下载链接】zeekZeek is a powerful network analysis framework that is much different from the typical IDS you may know.项目地址https://gitcode.com/gh_mirrors/ze/zeek点击查看免费下载本篇技术指南以 Zeek 仓库中doc/scripts/base/frameworks/sumstats/non-cluster.zeek.rst所对应的核心脚本 scripts/base/frameworks/sumstats/non-cluster.zeek 为主体深入讲解 Summary StatisticsSumStats框架在非集群单进程部署下的完整实现包括 epoch 周期的调度与分批处理、epoch_result / epoch_finished 回调触发时机、阈值threshold检测与跨阈值回调、手动 epoch 模式以及zeek_done()时的收尾逻辑。读完本文你将掌握 SumStats 框架三大核心概念Observation、Reducer、SumStat的落地实现细节并能在自己的 Zeek 脚本中正确编写基于 SumStats 的连接计数、扫描检测等实战脚本。一、背景为什么需要 SumStats 框架官方文档 doc/frameworks/sumstats.rst 开门见山地说明了框架的动机测量网络流量特征是 Zeek 脚本中最常见的任务之一。在简单的场景例如处理大小有限的 trace 文件下Zeek 自带的数据结构已经足够但真实部署面临两大困难集群化流量被多个 worker 进程同时抓取统计结果需要在各节点间汇总无界数据集流量永不停止内存中不能无限累积数据。SumStatsSummary Statistics框架正是为消费无界数据集并使其在大型集群和非集群 Zeek 部署中可实际测量而设计的机制。它在集群场景下通过 cluster.zeek 实现透明的跨节点聚合而在单进程场景下则依赖本文主角 non-cluster.zeek 完成 epoch 周期内的数据收集、阈值判断与结果回调。加载逻辑集群与非集群的二选一scripts/base/frameworks/sumstats/__load__.zeek是框架的入口其加载策略清晰体现了非集群实现的定位load ./main load ./plugins # The cluster framework must be loaded first. load base/frameworks/cluster # Load either the cluster support script or the non-cluster support script. if ( Cluster::is_enabled() ) load ./cluster else load ./non-cluster endif也就是说当且仅当集群框架未启用时non-cluster.zeek才会被加载。这也解释了为什么non-cluster.zeek中request_key()的函数注释写着This only needs to be implemented this way for cluster compatibility——它的存在是为了让非集群环境也能与集群版本的函数签名保持兼容见下文第六节。二、框架三大核心概念术语在深入non-cluster.zeek之前先明确官方文档定义的三段式处理流程这也是 scripts/base/frameworks/sumstats/main.zeek 中各个类型的直接来源概念官方定义对应源码类型Observation观测单个数据点属于某个任意命名的观测流observation stream拥有一个 Key 和实际观测值SumStats::Observationnum/dbl/str三选一、SumStats::KeyReducer归约器对观测流施加计算把无界观测集合归约为更小的表示按 Key 收集结果SumStats::Reducer含stream、apply、pred、normalize_keySumStat汇总统计最终定义一个或多个 Reducer 在一个时间间隔epoch内收集支持阈值与回调SumStats::SumStat含name、epoch、reducers、各类回调关键数据结构来自 main.zeekSumStats::Keystr: string optional与host: addr optional二选一或组合表示统计作用的对象例如发起连接的源 IP或 HTTP Host 头对应的非主机维度值。SumStats::Observation每次观测只提供一个字段——num: count、dbl: double、str: string。源码注释特别强调Only supply a single field at a time!。SumStats::Reducerstream是观测流标识符apply是需要执行的计算集合可选pred谓词按 Key 决定是否接受数据与normalize_key对 Key 做归一化/聚合。SumStats::ResultVal单个流在某个 Key 上的计算结果基础字段包括begin首条观测时间、end末条观测时间、num观测条数其余字段由计算插件追加如$sum、$unique。SumStats::Resulttable[string] of ResultVal按观测流标识符索引多个 Reducer 的结果。SumStats::ResultTabletable[Key] of Result按 Key 索引的结果总表——这正是non-cluster.zeek中process_epoch_result事件所处理的数据结构。SumStats::SumStat记录包含以下核心字段均来自 main.zeekname: stringSumStat 的任意名称用于后续引用epoch: intervalepoch 周期即数据被切断并触发epoch_result回调的间隔epoch 传0 secs表示手动 epoch 模式需自行调用SumStats::next_epoch结束周期reducers: set[Reducer]该 SumStat 关联的 Reducer 集合threshold_val/threshold/threshold_series阈值计算函数、单一阈值、阈值序列必须升序排列且后一个阈值只有在前一个被越过后才会检查threshold_crossed阈值被越过时的回调每个 epoch 内每个 Key 只触发第一次epoch_resultepoch 结束时按 Key 逐个回调epoch_finished整个 epoch 收集完成时的回调ts为收集开始时间。三、non-cluster.zeek 的整体工作流non-cluster.zeek在单进程环境下承担了main.zeek中两个内部全局函数/事件的实现data_added数据写入后的阈值检查入口与finish_epoch结束一个测量周期的事件。整体流程可概括为脚本调用SumStats::create()注册 SumStat由 main.zeek 实现若epoch ! 0secsmain.zeek 会schedule ss$epoch { SumStats::finish_epoch(ss) }定时触发finish_epoch事件到达后do_finish_epoch()从result_store取出该 SumStat 的全部结果若数据非空则触发SumStats::process_epoch_result事件把结果分批交给用户提供的epoch_result回调处理处理完成后reset(ss)清空结果与阈值状态并重新调度下一轮finish_epoch每条观测写入后data_added()调用check_thresholds()判断是否越界越界则触发threshold_crossed()。四、epoch 结果的分批处理process_epoch_resultnon-cluster.zeek中最核心的机制是process_epoch_result事件。由于单个 epoch 内可能积累大量 Key每个 Key 代表一个被统计对象一次事件处理中不可能也无必要同步遍历全部结果因此它采用分批 事件续排的方式event SumStats::process_epoch_result(ss: SumStat, now: time, data: ResultTable) { # TODO: is this the right processing group size? local i 50; local keys_to_delete: vector of SumStats::Key vector(); for ( key, res in data ) { ss$epoch_result(now, key, res); keys_to_delete key; if ( --i 0 ) break; } for ( idx in keys_to_delete ) delete data[keys_to_delete[idx]]; if ( |data| 0 ) # TODO: is this the right interval? schedule 0.01 secs { SumStats::process_epoch_result(ss, now, data) }; else if ( ss?$epoch_finished ) ss$epoch_finished(now); }几个值得注意的工程细节批大小为 50每次事件循环最多调用 50 次epoch_result回调避免单次事件占用过久的事件循环时间片源码中以 TODO 注释保留了对该数值合理性的讨论先收集后删除遍历过程中先把已处理的 Key 收集进keys_to_delete向量循环结束后统一delete避免在遍历表的同时修改表续排调度若表内仍有剩余 Key则schedule 0.01 secs重新触发同一事件继续处理把长任务切分成多个短事件保证事件循环对其他事件如新数据观测保持响应收尾回调当表被清空后若定义了epoch_finished回调则调用之标志整个 epoch 的数据已全部交付完毕。五、epoch 的结束与重置do_finish_epoch 与 finish_epochfinish_epoch事件是每个 epoch 周期的闹钟。do_finish_epoch()是它的实际实现体function do_finish_epoch(ss: SumStat) { if ( ss$name !in result_store || ! ss?$epoch_result ) return; local data result_store[ss$name]; local now network_time(); if ( zeek_is_terminating() ) { for ( key, val in data ) ss$epoch_result(now, key, val); if ( ss?$epoch_finished ) ss$epoch_finished(now); } else { if ( |data| 0 ) event SumStats::process_epoch_result(ss, now, copy(data)); else { if ( ss?$epoch_finished ) ss$epoch_finished(now); } } # We can reset here because we know that the reference to the # data will be maintained by the process_epoch_result event. reset(ss); if ( ss$epoch ! 0secs ) schedule ss$epoch { SumStats::finish_epoch(ss) }; }这里体现了两个重要的分支场景正常运行时若结果表非空则通过event SumStats::process_epoch_result(ss, now, copy(data))投递事件处理。注意传入的是copy(data)的副本而注释明确说明之所以可以放心在投递后立即reset(ss)是因为process_epoch_result事件持有那份副本的引用原始result_store可以安全清空重建。进程终止时zeek_is_terminating()Zeek 正在退出事件调度可能不再可靠执行因此改为同步for循环直接逐 Key 调用epoch_result回调随后调用epoch_finished确保统计数据在进程退出前完整输出。reset(ss)会清空result_store[ss$name]与threshold_tracker[ss$name]阈值追踪状态为下一个 epoch 提供干净起点。最后若epoch ! 0secs非手动模式重新调度下一轮finish_epoch形成周期性的统计节拍。finish_epoch事件本身还有一个终止保护的细节event SumStats::finish_epoch(ss: SumStat) { if ( zeek_is_terminating() ) return; # runs during zeek_done() instead do_finish_epoch(ss); }即进程退出时跳过定时触发的finish_epoch改由下面的zeek_done()统一收尾避免重复处理。六、zeek_done() 收尾与非手动 epoch 的强制结算non-cluster.zeek对退出阶段的处理非常考究# Run non-manual SumStats entries as late as possible, but a bit # earlier than a users zeek_done() handler in case they end up # doing something curious in zeek_done(). event zeek_done() priority10 { for ( name, ss in stats_store ) { if ( ss$epoch ! 0sec ) # skip SumStats with manual epochs. do_finish_epoch(ss); } }要点如下给zeek_done()设置priority10使其晚于默认优先级的事件执行尽可能覆盖完整运行期同时又早于用户自己的zeek_done()处理器源码注释的担心是用户可能在zeek_done()里做出人意料的事情保证统计结算先行完成遍历stats_store中所有注册的 SumStat跳过手动 epoch 模式epoch 0sec的条目——手动模式由用户自行控制结束时机框架不代劳对非手动 SumStat 调用do_finish_epoch(ss)此时zeek_is_terminating()为真走同步回调分支把最后一个 epoch 的残留数据全部交付给epoch_result/epoch_finished。request_key为集群兼容而存在的动态查询function request_key(ss_name: string, key: Key): Result { # This only needs to be implemented this way for cluster compatibility. return when [ss_name, key] ( T ) { if ( ss_name in result_store key in result_store[ss_name] ) return result_store[ss_name][key]; else return table(); } }注释直言该实现仅为了集群兼容性。它被设计为可在when语句中异步调用的函数直接查本地result_store并返回结果集群版本cluster.zeek则通过网络向各 worker 广播查询再聚合。官方文档 main.zeek 提醒request_key应谨慎使用不能替代SumStat记录自带的回调机制。七、阈值检测data_added → check_thresholds → threshold_crossed观测数据写入SumStats::observe()实现在 main.zeek会最终调用data_added。非集群实现把它定义为最直接的每次写入即检查function data_added(ss: SumStat, key: Key, result: Result) { if ( check_thresholds(ss, key, result, 1.0) ) threshold_crossed(ss, key, result); }对比集群版本 cluster.zeek集群的data_added使用cluster_request_global_view_percent默认0.2做预判达到全量阈值 20% 时向 manager 发中间更新请求实现 epoch 中途的提前告警非集群版本则直接用modify_pct 1.0做完整阈值判断无需网络协调。check_thresholdsmain.zeek 实现的判定逻辑先补齐result中缺失的 Reducer 结果用init_resultval初始化空值保证threshold_val编写时能安全访问所有流调用ss$threshold_val(key, result)得到观测值watch单一阈值t_index 0该 Key 尚未越过阈值且watch ss$threshold即返回真阈值序列|ss$threshold_series| t_index还有未越过的阈值且watch ss$threshold_series[t_index]即返回真。threshold_crossedmain.zeek 实现在被调用时会increment_threshold_tracker(ss$name, key)记录该 Key 已越过当前阈值序列模式下数值递增指向下一个待越阈值同样补齐缺失的 ResultVal调用用户提供的ss$threshold_crossed(key, result)回调。八、实战示例从连接到扫描检测官方文档 doc/frameworks/sumstats.rst 提供了两个可运行的完整脚本直接继承如下。示例一统计连接数完整脚本见 doc/frameworks/sumstats-countconns.zeekload base/frameworks/sumstats event connection_established(c: connection) { # Make an observation! # This observation is global so the key is empty. # Each established connection counts as one so the observation is always 1. SumStats::observe(conn established, SumStats::Key(), SumStats::Observation($num1)); } event zeek_init() { # Create the reducer. # The reducer attaches to the conn established observation stream # and uses the summing calculation on the observations. local r1 SumStats::Reducer($streamconn established, $applyset(SumStats::SUM)); # Create the final sumstat. # We give it an arbitrary name and make it collect data every minute. # The reducer is then attached and a $epoch_result callback is given # to finally do something with the data collected. SumStats::create([$name counting connections, $epoch 1min, $reducers set(r1), $epoch_result(ts: time, key: SumStats::Key, result: SumStats::Result) { # This is the body of the callback that is called when a single # result has been collected. We are just printing the total number # of connections that were seen. The $sum field is provided as a # double type value so we need to use %f as the format specifier. print fmt(Number of connections established: %.0f, result[conn established]$sum); }]); }在 Zeek 测试套件的样例 PCAP 上运行官方文档给出的输出为$ zeek -r workshop_2011_browse.trace sumstats-countconns.zeek Number of connections established: 6注意点这里Key()为空全局统计每个 epoch1 分钟结束时process_epoch_result会为唯一 Key 触发一次epoch_result通过result[conn established]$sum读取 SUM 插件累积的和。由于$sum是 double 类型格式化需用%f。示例二玩具扫描检测阈值功能演示完整脚本见 doc/frameworks/sumstats-toy-scan.zeekload base/frameworks/sumstats # We use the connection_attempt event to limit our observations to those # which were attempted and not successful. event connection_attempt(c: connection) { # Make an observation! # This observation is about the host attempting the connection. # Each established connection counts as one so the observation is always 1. SumStats::observe(conn attempted, SumStats::Key($hostc$id$orig_h), SumStats::Observation($num1)); } event zeek_init() { # Create the reducer. # The reducer attaches to the conn attempted observation stream # and uses the summing calculation on the observations. Keep # in mind that there will be one result per key (connection originator). local r1 SumStats::Reducer($streamconn attempted, $applyset(SumStats::SUM)); # Create the final sumstat. # This is slightly different from the last example since were providing # a callback to calculate a value to check against the threshold with # $threshold_val. The actual threshold itself is provided with $threshold. # Another callback is provided for when a key crosses the threshold. SumStats::create([$name finding scanners, $epoch 5min, $reducers set(r1), # Provide a threshold. $threshold 5.0, # Provide a callback to calculate a value from the result # to check against the threshold field. $threshold_val(key: SumStats::Key, result: SumStats::Result) { return result[conn attempted]$sum; }, # Provide a callback for when a key crosses the threshold. $threshold_crossed(key: SumStats::Key, result: SumStats::Result) { print fmt(%s attempted %.0f or more connections, key$host, result[conn attempted]$sum); }]); }对包含 nmap 扫描主机的 PCAP 运行官方文档给出的输出为$ zeek -r nmap-vsn.trace sumstats-toy-scan.zeek 192.168.1.71 attempted 5 or more connections这个示例完整展示了非集群阈值链路每次connection_attempt观测写入后data_added立即执行check_thresholds(ss, key, result, 1.0)当某源 IP 的SUM值达到$threshold 5.0时threshold_crossed回调被触发并打印告警。文档也提醒这只是演示阈值机制的教学玩具并非生产级扫描检测方案。九、计算插件体系SumStats 的可扩展计算non-cluster.zeek本身不含具体计算逻辑——所有的计算Calculation都通过插件机制注册见 scripts/base/frameworks/sumstats/plugins/ 目录插件文件计算说明sum.zeekSUM数值累加字符串观测按 1.0 计数average.zeekAVG平均值variance.zeekVARIANCE方差std-dev.zeekSTD_DEV标准差unique.zeekUNIQUE唯一值计数支持unique_max上限hll_unique.zeekHLL_UNIQUE基于 HyperLogLog 的基数估计max.zeek / min.zeekMAX/MIN最大值 / 最小值last.zeekLAST最近一个观测值sample.zeekSAMPLE采样topk.zeekTOPKTop-K 统计以 sum.zeek 为例可以看清插件的三个挂钩点init_resultval_hook初始化 ResultVal 时给$sum字段赋默认值 0register_observe_plugins把SUM计算注册为rv$sum val的观测函数compose_resultvals_hook集群合并结果时把两个 ResultVal 的$sum相加compose_resultvals在 main.zeek 中调用该挂钩。这条注册 → 观测 → 合并的链路正是非集群与集群共用同一套计算语义的底层保证非集群模式不经历网络合并但插件接口完全一致。十、内部状态存储与生命周期理解non-cluster.zeek还需要知道它操作的四张全局表均定义于 main.zeekstats_store: table[string] of SumStat按名称索引所有已注册的 SumStatreducer_store: table[string] of set[Reducer]按观测流标识符索引 Reducer 集合observe()据此找到应处理该观测的所有 Reducerresult_store: table[string] of ResultTable按 SumStat 名称索引的当前 epoch 结果总表——do_finish_epoch从这里取数process_epoch_result逐批消费reset()清空重建threshold_tracker: table[string] of table[Key] of count记录每个 SumStat 下每个 Key 已越过的阈值序号单一阈值时 0 表示未越过序列时表示序列偏移量。SumStats::create()在 main.zeek 中还会对每个 Reducer 展开计算依赖add_calc_deps并校验给出了阈值但没有threshold_val函数的错误用法Reporter::error。这些状态表的读写与non-cluster.zeek的事件循环共同构成了单进程模式下的完整生命周期注册 → 定时 epoch → 分批回调 → 阈值检测 → 重置 → 再调度 → 退出收尾。十一、总结non-cluster.zeek是 Zeek SumStats 框架在单进程部署下的完整运行引擎它的设计处处为无界数据 事件驱动服务分批处理每批 50 个 Key 0.01 秒续排防止长任务阻塞事件循环运行时事件化、退出时同步化的双路径do_finish_epoch保证数据不丢失zeek_done() priority10在用户收尾代码之前完成所有非手动 SumStat 的最终结算data_added的即时阈值检查让扫描类检测能在 epoch 中途快速响应request_key的本地实现保持与集群版 API 兼容。掌握这份脚本你就能透彻理解 SumStats 框架为何能在小型单进程 Zeek 实例和大型多 worker 集群上同样工作——集群的复杂协调逻辑cluster.zeek被隔离在if ( Cluster::is_enabled() )之后而非集群路径则保持最简、最高效的实现二者共享 main.zeek 定义的数据模型与插件计算体系。赞分享网络安全网络IDS【免费下载链接】zeekZeek is a powerful network analysis framework that is much different from the typical IDS you may know.项目地址https://gitcode.com/gh_mirrors/ze/zeek点击查看免费下载相关推荐Zeek SumStats 框架详解用观察流、Reducer 与 Epoch 构建可伸缩的网络统计与阈值检测Zeek SumStats 框架详解用观察流、Reducer 与 Epoch 构建可伸缩的网络统计与阈值检测 导读 本文以 Zeek 的 Summary St网络安全网络IDSZeek SumStats 框架实战指南面向集群与无界流量数据的流式统计与阈值检测Zeek SumStats 框架实战指南面向集群与无界流量数据的流式统计与阈值检测 本篇技术指南围绕 Zeek 内置的 SumStatsSummary St网络安全网络IDSbigdata_analyse 数据处理性能优化10个提升效率的技巧bigdata_analyse 数据处理性能优化10个提升效率的技巧 在大数据分析项目中数据处理性能优化是提升整体效率的关键环节。无论你是在处理百万级的用户网络安全网络IDS上一篇capa 版本演进与技术全景从 v1.0 到 master 的可执行文件能力识别引擎发展脉络下一篇Wand-Enhancer终极指南免费解锁Wand游戏修改器完整功能创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表