ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

H2O-3 PCA 的 `impute_missing` 参数实战指南:用列均值填补缺失值,避免训练行被静默丢弃

H2O-3 PCA 的 `impute_missing` 参数实战指南:用列均值填补缺失值,避免训练行被静默丢弃 机器学习深度学习AutoML大数据后端【免费下载链接】h2o-3H2O is an Open Source, Distributed, Fast Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.项目地址https://gitcode.com/gh_mirrors/h2/h2o-3点击查看免费下载impute_missing是 H2O-3 中 PCA主成分分析h2o.prcomp专用的布尔参数用于控制训练数据中的 NA/缺失值处理方式。本文基于 impute_missing.rst 官方参数文档并结合 H2O-3 仓库中 PCA 的 Java 实现源码完整讲解该参数的作用机制、默认行为、底层实现细节以及 R 与 Python 双语言的完整实战示例。读完本文你将掌握如何判断 PCA 训练时数据行是否被缺失值剔除、如何用impute_missingTRUE开启列均值填补、如何解读开启前后主成分结果的变化以及该参数与pca_method各实现路径之间的协作关系。参数速览项目说明参数名impute_missing可用算法PCA主成分分析是否为超参数否Hyperparameter: no不参与网格搜索/调参迭代默认值False参数类型布尔值Boolean作用用每列的均值column mean填补各列的缺失条目而不是删除含缺失值的行在 R 客户端中通过h2o.prcomp(impute_missing TRUE)传入在 Python 客户端中通过H2OPrincipalComponentAnalysisEstimator(impute_missingTrue)传入。Python 客户端的参数说明在 pca.py 中定义为 Whether to impute missing entries with the column mean默认值同样是False见 pca.py。为什么需要这个参数缺失值导致的行剔除问题PCA 本质上是对协方差/相关结构进行特征分解绝大多数实现都要求输入矩阵中不含缺失值。H2O-3 的默认行为impute_missingFalse是训练前检测训练集中是否存在 NA若存在 NA则通过(na.omit ...)操作整行删除所有包含缺失值的样本因此实际参与分解的行数nobs可能远小于原始数据行数当缺失行过多、导致有效行数小于主成分数k时训练直接报错。这在源码中体现得非常清晰。在 PCA.java 中当impute_missing为假且数据含 NA 时训练流程会先发出警告然后执行na.omit删行if (!_parms._impute_missing) { // added warning to user per request from Nidhi _job.warn(_train: Dataset used may contain fewer number of rows due to removal of rows with NA/missing values. If this is not desirable, set impute_missing argument in pca call to TRUE/True/true/... depending on the client language.); } if ((!_parms._impute_missing) tranRebalanced.hasNAs()) { // remove NAs rows ... _train Rapids.exec(String.format((na.omit %s), tranRebalanced._key)).getFrame(); // remove NA rows ... }如果删行后有效样本数nobs小于kPCA.java 会直接抛出参数校验错误并给出三条补救建议if((model._output._nobs 0) || (model._output._nobs _parms._k )) { error(_train, Number of row in _train is less than k. Consider setting impute_missing TRUE or using pca_method GLRM instead or reducing the value of parameter k.); }也就是说impute_missingTRUE正是解决缺失值导致有效行数不足 / 有效样本量缩水这类问题的最直接手段。开启后发生了什么列均值填补的底层机制参数在模型中的定义impute_missing首先作为 PCA 模型参数被定义在 PCAModel.javapublic boolean _impute_missing false; // Should missing numeric values be imputed with the column mean?注释明确该参数决定缺失的数值是否用列均值填补。同时在 REST 层PCAV3.java 将其注册为公开的 schema 属性见该文件第 27 行参数名注册与第 73 行字段声明因此所有客户端R/Python/Flow/REST API都能以统一方式读写它。数据准备阶段DataInfo 的 skipMissing / imputeMissing训练真正开始时PCA 会基于训练集构造DataInfo对象其中两个布尔位与impute_missing直接挂钩PCA.javadinfo new DataInfo(_train, _valid, 0, _parms._use_all_factor_levels, _parms._transform, DataInfo.TransformType.NONE, /* skipMissing */ !_parms._impute_missing, /* imputeMissing */ _parms._impute_missing, /* missingBucket */ false, /* weights */ false, /* offset */ false, /* fold */ false, /* intercept */ false);skipMissing !impute_missing为真时跳过含缺失值的行imputeMissing impute_missing为真时在特征化阶段用列均值自动填补缺失的数值。均值与标准差的先算后用一个容易被忽视的细节是均值和标准差的归属问题如果先删行再计算均值那么均值只代表无缺失子集的均值与原始数据分布有偏差。源码对此做了专门处理PCA.javaif (!_parms._impute_missing tranRebalanced.hasNAs()) { // fixed the std and mean of dinfo to that of the frame before removing NA rows dinfo._normMul tinfo._normMul; dinfo._numMeans tinfo._numMeans; dinfo._numNAFill dinfo._numMeans; // NAs will be imputed with means dinfo._normSub tinfo._normSub; }这里先用未删行的tinfo计算好均值_numMeans与标准差再让_numNAFill指向这些均值保证缺失值是用全量数据的列均值填补的。也就是说即使关闭impute_missing走删行路径标准化参数仍来自删除前的完整数据避免系统性偏差。内存检查的时机差异另外PCA.java 显示内存占用检查checkMemoryFootPrint的执行时机也与缺失值有关当训练集不含 NA或开启了impute_missing时会在初始化阶段提前执行内存检查而在删行路径下会等na.omit完成后再检查一次见第 289 行。这是为宽数据集wide dataset场景准备的防御性设计。不同 pca_method 对 impute_missing 的处理差异H2O-3 PCA 支持四种pca_methodimpute_missing在每种路径下的传递方式并不相同从源码可以明确区分GramSVD默认直接基于DataInfo计算 Gram 矩阵AA/n后做 SVD。缺失值处理完全由上述skipMissing/imputeMissing逻辑控制impute_missing参数直接生效。Power / Randomized幂迭代 / 随机化这两种方法在内部委托给 SVD 模型求解PCA 会把_impute_missing原样传给 SVD 参数PCA.javaparms._impute_missing _parms._impute_missing;对应的 SVD 实现SVD.java与 PCA 采用完全一致的na.omit 均值填补逻辑甚至连NAs will be imputed with means的处理都是同构的。GLRM从 PCA.java 的代码结构看GLRM 分支构造GLRMParameters时并未传递impute_missing。这是因为 GLRM 通过矩阵因子化低秩分解天然支持缺失数据——缺失条目在损失函数中被跳过不需要预先删行或均值填补。因此可以推断当数据缺失严重、均值填补过于粗糙时直接使用pca_methodGLRM是比impute_missingTRUE更温和的替代方案这与源码错误提示中Consider setting impute_missing TRUE or using pca_method GLRM的建议完全一致。R 实战示例Birds 数据集的前后对比以下是官方文档提供的 R 示例使用 H2O 公开测试数据birds.csv完整展示开启与关闭impute_missing两种场景library(h2o) h2o.init() # Load the Birds dataset birds - h2o.importFile(https://s3.amazonaws.com/h2o-public-test-data/smalldata/pca_test/birds.csv) # Train with impute_missing enabled birds_pca - h2o.prcomp(training_frame birds, transform STANDARDIZE, k 3, pca_method Power, use_all_factor_levels TRUE, impute_missing TRUE) # View the importance of components birds_pcamodel$importance Importance of components: pc1 pc2 pc3 Standard deviation 1.496991 1.351000 1.014182 Proportion of Variance 0.289987 0.236184 0.133098 Cumulative Proportion 0.289987 0.526171 0.659269 # View the eigenvectors birds_pcamodel$eigenvectors Rotation: pc1 pc2 pc3 patch.Ref1a 0.007207 0.007449 0.001161 patch.Ref1b -0.003090 0.011257 -0.001066 patch.Ref1c 0.002962 0.008850 -0.000264 patch.Ref1d -0.001295 0.011003 0.000501 patch.Ref1e 0.006559 0.006904 -0.001206 --- pc1 pc2 pc3 S 0.463591 -0.053410 0.184799 year -0.055934 0.009691 -0.968635 area 0.533375 -0.289381 -0.130338 log.area. 0.583966 -0.262287 -0.089582 ENN -0.270615 -0.573900 0.038835 log.ENN. -0.231368 -0.640231 0.026325 # Train again without imputing missing values birds2_pca - h2o.prcomp(training_frame birds, transform STANDARDIZE, k 3, pca_method Power, use_all_factor_levels TRUE, impute_missing FALSE) Warning message: In doTryCatch(return(expr), name, parentenv, handler) : _train: Dataset used may contain fewer number of rows due to removal of rows with NA/missing values. If this is not desirable, set impute_missing argument in pca call to TRUE/True/true/... depending on the client language. # View the importance of components birds2_pcamodel$importance Importance of components: pc1 pc2 pc3 Standard deviation 1.546397 1.348276 1.055239 Proportion of Variance 0.300269 0.228258 0.139820 Cumulative Proportion 0.300269 0.528527 0.668347 # View the eigenvectors birds2_pcamodel$eigenvectors Rotation: pc1 pc2 pc3 patch.Ref1a 0.009848 -0.005947 -0.001061 patch.Ref1b -0.001628 -0.014739 -0.001007 patch.Ref1c 0.004994 -0.009486 -0.000523 patch.Ref1d 0.000117 -0.004400 -0.004917 patch.Ref1e 0.003627 -0.001467 -0.004268 --- pc1 pc2 pc3 S 0.515048 0.226915 -0.123136 year -0.066269 -0.069526 0.971250 area 0.414050 0.344332 0.149339 log.area. 0.497313 0.363609 0.131261 ENN -0.390235 0.545631 -0.007944 log.ENN. -0.345665 0.562834 -0.002092对比两组结果可以观察到警告信息当impute_missingFALSE时R 控制台会收到上述警告明确提示数据集可能因移除 NA/缺失行而包含更少的行这正是 PCA.java 中_job.warn(...)的原文——R 客户端把它以Warning message形式透出。重要性指标变化开启填补后Standard deviation1.497/1.351/1.014低于删行版本1.546/1.348/1.055累计方差贡献率也从 0.668 降到 0.659。这说明两种预处理方式对应不同的有效样本量nobs与 Gram 矩阵结果不能简单等同。特征向量变化Rotation表中的系数如S、year、area等数值列在两种模式下差异明显且year对 pc3 的贡献方向甚至发生反转-0.968635vs0.971250。因此当数据含缺失值且业务上要求全量样本参与分析时务必明确选择并记录impute_missing的设置。Python 实战示例官方文档同样提供了 Python 版本使用H2OPrincipalComponentAnalysisEstimator估计器import(h2o) h2o.init() from h2o.estimators.pca import H2OPrincipalComponentAnalysisEstimator # Load the Birds dataset birds h2o.import_file(https://s3.amazonaws.com/h2o-public-test-data/smalldata/pca_test/birds.csv) # Train with impute_missing enabled birds.pca H2OPrincipalComponentAnalysisEstimator(k 3, transform STANDARDIZE, pca_methodPower, use_all_factor_levelsTrue, impute_missingTrue) birds.pca.train(xlist(range(4)), training_framebirds) # View the importance of components birds.pca.varimp(use_pandasFalse) [(uStandard deviation, 1.0505993078459912, 0.8950182545325247, 0.5587566783073901), (uProportion of Variance, 0.28699613488673914, 0.20828865401845226, 0.08117966990084355), (uCumulative Proportion, 0.28699613488673914, 0.4952847889051914, 0.5764644588060349)] # View the eigenvectors birds.pca.rotation() Rotation: pc1 pc2 pc3 ---------------- ------------------ ----------------- ---------------- patch.Ref1a 0.00732398141913 -0.0141576160836 0.0294419461081 patch.Ref1b -0.00482860843905 0.00867426840498 0.0330778190153 patch.Ref1c 0.00124768649004 -0.00274167383932 0.0312598825617 patch.Ref1d -0.000370181920761 0.000297923901103 0.0317439245635 patch.Ref1e 0.00223394447742 -0.00459462277502 0.0309648089406 --- --- --- --- landscape.Bauxite -0.0638494513759 0.136728811833 0.118858152002 landscape.Forest 0.0378085502606 -0.0833578672691 0.969316569884 landscape.Urban -0.0545759062856 0.111309410422 0.0354475756223 S 0.564501605704 -0.767095710638 -0.0466832766991 year -0.814596906726 -0.577331674836 -0.0101626722479 # See the whole table with table.as_data_frame() # Train again without imputing missing values birds2 h2o.import_file(https://s3.amazonaws.com/h2o-public-test-data/smalldata/pca_test/birds.csv) birds2.pca H2OPrincipalComponentAnalysisEstimator(k 3, transform STANDARDIZE, pca_methodPower, use_all_factor_levelsTrue, impute_missingFalse) birds2.pca.train(xlist(range(4)), training_framebirds2) # View the importance of components birds2.pca.varimp(use_pandasFalse) [(uStandard deviation, 1.1238486420242524, 0.949554306091356, 0.534896629598228), (uProportion of Variance, 0.3080623966646966, 0.21991895069672512, 0.06978510918460899), (uCumulative Proportion, 0.3080623966646966, 0.5279813473614217, 0.5977664565460307)] # View the eigenvectors birds2.pca.rotation() Rotation: pc1 pc2 pc3 ---------------- ----------------- ----------------- ----------------- patch.Ref1a 0.00898674970716 0.0133755203176 0.0386887315027 patch.Ref1b -0.00583910665399 -0.00850852817775 0.0403921679996 patch.Ref1c 0.00157382152659 0.00243349606991 0.0395404497512 patch.Ref1d 0.00205431391489 -0.00464763108225 0.0130225730145 patch.Ref1e 0.00521157104675 9.98792622547e-07 0.0126676559841 --- --- --- --- landscape.Bauxite -0.0927064158093 -0.0985077050027 0.312254932996 landscape.Forest 0.049803344754 0.0606680349608 0.928822693132 landscape.Urban -0.0671561320808 -0.108679950396 0.033639706807 S 0.661206203315 0.69412159594 -0.0166591571667 year -0.727793152951 0.684904477663 -0.00409291536614 # See the whole table with table.as_data_frame()Python 侧的使用要点参数既可以在构造估计器时传入H2OPrincipalComponentAnalysisEstimator(impute_missingTrue)也可以在训练前通过属性赋值设置estimator.impute_missing True属性 setter 位于 pca.py并会做布尔类型断言官方示例中xlist(range(4))表示取前 4 列作为预测特征实际使用时请按自己的特征列名如x[S, year, area]显式指定varimp(use_pandasFalse)返回三元组列表Standard deviation / Proportion of Variance / Cumulative Proportionrotation()返回特征向量表可通过table.as_data_frame()转为 DataFrame 查看完整结果。使用建议与注意事项理解均值填补的代价列均值填补会人为压缩各列的方差、扭曲变量间的相关结构尤其当缺失比例很高时主成分方向可能被系统性拉偏。它是保留行数与保留原始分布之间的权衡impute_missingTRUE保住样本量但引入插补偏差impute_missingFALSE保住原始观测但牺牲样本量与代表性。缺失严重的场景优先考虑 GLRM如源码 PCA.java 的提示所示缺失值占比过高时pca_methodGLRM是比均值填补更稳健的选择因为 GLRM 在因子化过程中天然跳过缺失条目。注意nobs k的训练错误若数据缺失导致删行后有效行数不足k训练会失败。此时应开启impute_missingTRUE、改用 GLRM或调低k值。该参数不是超参数impute_missing标记为 Hyperparameter: no不适合放入网格搜索H2OGridSearch的hyper_params字典中参与调参它应当作为数据预处理决策在训练前一次性确定。与其他预处理参数的配合本参数与transformSTANDARDIZE/NORMALIZE等、use_all_factor_levels独立生效文档明确列出其 Related Parameters 为 None即不存在强耦合的关联参数。客户端行为一致性无论 R、Python 还是 Flow该参数都以同名布尔值暴露不同语言接受TRUE/True/true等对应写法这是 H2O 客户端层的通用约定。深入阅读官方参数文档h2o-docs/src/product/data-science/algo-params/impute_missing.rstPCA 训练主流程与na.omit/均值填补实现h2o-algos/src/main/java/hex/pca/PCA.java模型参数定义默认值false与注释说明h2o-algos/src/main/java/hex/pca/PCAModel.javaREST schema 注册h2o-algos/src/main/java/hex/schemas/PCAV3.java同构逻辑的 SVD 实现Power/Randomized 路径委托对象h2o-algos/src/main/java/hex/svd/SVD.javaPython 客户端参数定义与属性访问器h2o-py/h2o/estimators/pca.py赞分享机器学习深度学习AutoML大数据后端【免费下载链接】h2o-3H2O is an Open Source, Distributed, Fast Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.项目地址https://gitcode.com/gh_mirrors/h2/h2o-3点击查看免费下载相关推荐告别参数缺失go-swagger API默认值3步配置指南告别参数缺失go swagger API默认值3步配置指南 你是否曾因API参数缺失导致服务崩溃是否还在手动校验每个请求的必填项本文将通过3个简单步骤教代码生成开发工具后端API设计如何用WinUtil快速完成Windows系统优化与软件管理如何用WinUtil快速完成Windows系统优化与软件管理 你是否曾经为新安装的Windows系统感到头疼需要一个个安装常用软件、调整系统设置、优化性能这桌面应用运维用 Kornia AverageMeter 构建训练循环的运行均值监控kornia.metrics.monitoring 实战指南用 Kornia AverageMeter 构建训练循环的运行均值监控kornia.metrics.monitoring 实战指南 kornia.metric计算机视觉人工智能深度学习图像处理上一篇30分钟零门槛用create-react-native-app开发蓝牙低功耗(BLE)应用下一篇CANN opbase 算子开发指南aclTensor::SetFloatData 主机侧 Tensor 浮点数据写入详解创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表