Research Article | | Peer-Reviewed

Feature Selection in Over-dispersed Binary and Count Data Models Using Penalized Optimal Estimating Functions

Received: 12 August 2026     Accepted: 5 September 2026     Published: 24 September 2026
Views:       Downloads:
Abstract

Generalized linear models (GLMs) remain a core class of supervised machine learning models for binary and count responses, with feature selection commonly carried out through penalized likelihood or quasi-likelihood methods. This paper develops a feature-selection framework based on penalized Optimal Estimating Functions (OEFs), which attain Godambe optimality within a class of unbiased estimating functions and incorporate higher-order moment information without requiring full likelihood specification. Ridge, LASSO, Adaptive LASSO and SCAD penalties are introduced at the regression estimating-equation level, while dispersion is estimated jointly through an unpenalized OEF. Hyperparameters are selected using cross-validated estimating-function loss with prespecified edge and stability rules. Monte Carlo simulations with 500 replications examine Beta-Binomial and Negative-Binomial regression under moderate and high overdispersion and sparse and moderately dense signals. The results do not show uniform superiority of penalization. The unpenalized OEF generally provides the strongest coefficient coverage and RMSE benchmark, whereas the penalized OEFs provide sparse feature selection with model-dependent trade-offs between sensitivity, specificity and interval calibration. Simple post-selection refitting reduces shrinkage bias in some settings but does not restore nominal coverage because selection uncertainty remains. Penalized OEF is therefore presented as a competitive alternative when sparse selection and joint mean-dispersion estimation are both required, rather than as a uniformly better estimator.

Published in American Journal of Theoretical and Applied Statistics (Volume 15, Issue 5)
DOI 10.11648/j.ajtas.20261505.17
Page(s) 276-302
Creative Commons

This is an Open Access article, distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution and reproduction in any medium or format, provided the original work is properly cited.

Copyright

Copyright © The Author(s), 2026. Published by Science Publishing Group

Keywords

Feature Selection, Penalized Estimating Functions, Over-dispersion, Beta-Binomial, Negative Binomial, Godambe Information, Post-selection Inference, Stability Selection

References
[1] Hoerl, A. E., Kennard, R. W. Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1), 55-67 1970.
[2] Tibshirani, R. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58(1), 267-288 1996.
[3] Zou, H. The adaptive lasso and its oracle properties. Journal of the American Statistical Association, 101(476), 1418-1429 2006.
[4] Fan, J., Li, R. Variable selection via nonconcave penalized likelihood and its oracle properties. Journal of the American Statistical Association, 96(456), 1348-1360 2001.
[5] McCullagh, P., Nelder, J. A. Generalized linear models (2nd ed.). Chapman & Hall 1989.
[6] Dean, C. B., Lundy, E. R. Overdispersion. In Wiley StatsRef: Statistics reference online. John Wiley & Sons 2016.
[7] Wedderburn, R. W. M. Quasi-likelihood functions, generalized linear models, and the Gauss-Newton method. Biometrika, 61(3), 439-447 1974.
[8] Godambe, V. P., Thompson, M. E. An extension of quasi-likelihood estimation. Journal of Statistical Planning and Inference, 22(2), 137-152 1989.
[9] Desmond, A. F. Optimal estimating functions, quasi-likelihood and statistical modelling. Journal of Statistical Planning and Inference, 60(1), 77-104 1997.
[10] Johnson, B. A., Lin, D. Y., Zeng, D. Penalized estimating functions and variable selection in semiparametric regression models. Journal of the American Statistical Association, 103(482), 672-680 2008.
[11] Godambe, V. P. The foundations of finite sample estimation in stochastic processes. Biometrika, 72(2), 419-428 1985.
[12] Najera-Zuloaga, J., Lee, D.-J., Arostegui, I. Comparison of beta-binomial regression model approaches to analyze health-related quality of life data. Statistical Methods in Medical Research, 27(10), 2989-3009 2018.
[13] Lehmann, B. D., Archer, K. J. Penalized negative binomial models for modeling an overdispersed count outcome with a high-dimensional predictor space: Application predicting micronuclei frequency. PLOS ONE, 14(1), Article e0209923 2019.
[14] Lindén, A., Mäntyniemi, S. Using the negative binomial distribution to model overdispersion in ecological count data. Ecology, 92(7), 1414-1421 2011.
[15] Cameron, A. C., Trivedi, P. K. Regression analysis of count data (2nd ed.). Cambridge University Press 2013.
[16] Suryadi, F., Jonathan, S., Jonatan, K., Ohyver, M. Handling overdispersion in Poisson regression using negative binomial regression for poverty case in West Java. Procedia Computer Science, 216, 517-523 2023.
[17] Heyde, C. C. Quasi-likelihood and its application: A general approach to optimal parameter estimation. New York, NY: Springer 1997.
[18] Hastie, T., Tibshirani, R., Friedman, J. The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer 2009.
[19] Bühlmann, P., van de Geer, S. Statistics for high-dimensional data: Methods, theory and applications. Springer 2011.
[20] Claeskens, G., Hjort, N. L. Model selection and model averaging. Cambridge University Press 2008.
[21] Stone, M. Cross-validatory choice and assessment of statistical predictions. Journal of the Royal Statistical Society: Series B (Methodological), 36(2), 111-147 1974.
[22] Komiyama, J., Maehara, T. Cross validation based model selection via generalized method of moments. arXiv 2018.
[23] Wang, F., Mukherjee, S., Richardson, S., Hill, S. M. High-dimensional regression in practice: An empirical study of finite-sample prediction, variable selection and ranking. Statistics and Computing, 30(3), 697-719 2020.
[24] Breiman, L., Friedman, J. H., Olshen, R. A., Stone, C. J. Classification and regression trees. Wadsworth 1984.
[25] Chen, Y., Yang, Y. The one standard error rule for model selection: Does it work?. Stats, 4(4), 868-892 2021.
[26] Yates, L. A., Aandahl, Z., Richards, S. A., Brook, B. W. Cross validation for model selection: A review with examples from ecology. Ecological Monographs, 92(1), Article e1557 2022.
[27] Donoho, D. L., Johnstone, I. M. Ideal spatial adaptation by wavelet shrinkage. Biometrika, 81(3), 425-455 1994.
[28] Sutradhar, B. C., Das, K. On joint estimation of regression and overdispersion parameters in generalized linear models for longitudinal data. Journal of Multivariate Analysis, 58(1), 90-108 1996.
[29] Paul, S. R., Islam, A. S. Joint estimation of the mean and dispersion parameters in the analysis of proportions. Canadian Journal of Statistics, 26(1), 83-94 1998.
[30] Harrison, X. A. Using observation-level random effects to model overdispersion in count data in ecology and evolution. PeerJ, 2, e616 2014.
[31] Zhang, C.-H. Nearly unbiased variable selection under minimax concave penalty. The Annals of Statistics, 38(2), 894-942 2010.
[32] Breheny, P., Huang, J. Coordinate descent algorithms for nonconvex penalized regression, with applications to biological feature selection. The Annals of Applied Statistics, 5(1), 232-253 2011.
[33] Meinshausen, N., Bühlmann, P. Stability selection. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 72(4), 417-473 2010.
[34] Ternès, N., Rotolo, F., Michiels, S. Empirical extensions of the lasso penalty to reduce the false discovery rate in high-dimensional Cox regression models. Statistics in Medicine, 35(15), 2561-2573 2016.
[35] Godambe, V. P. Estimating functions. Oxford University Press 1991.
[36] Riley, R. D., Snell, K. I. E., Martin, G. P., Whittle, R., Archer, L., Sperrin, M., Collins, G. S. Penalization and shrinkage methods produced unreliable clinical prediction models especially when sample size was small. Journal of Clinical Epidemiology, 132, 88-96 2020.
[37] van de Wiel, M. A., Leday, G. G. R., Heymans, M. W., van Zwet, E. W., Zwinderman, A. H., Hoogland, J. Alternatives to default shrinkage methods can improve prediction accuracy, calibration, and coverage: A methods comparison study. Statistical Methods in Medical Research, 34(7), 1342-1355 2025.
[38] Leeb, H., Pötscher, B. M. Can one estimate the conditional distribution of post-model-selection estimators?. The Annals of Statistics, 34(5), 2554-2591 2006.
[39] Berk, R., Brown, L., Buja, A., Zhang, K., Zhao, L. Valid post-selection inference. The Annals of Statistics, 41(2), 802-837 2013.
[40] Tibshirani, R. J., Taylor, J., Lockhart, R., Tibshirani, R. Exact post-selection inference for sequential regression procedures. Journal of the American Statistical Association, 111(514), 600-620 2016.
[41] Efron, B., Hastie, T. Computer age statistical inference: Algorithms, evidence, and data science. Cambridge University Press 2016.
[42] Hilbe, J. M. Negative binomial regression (2nd ed.). Cambridge University Press 2011.
[43] Kenne Pagui, E. C., Salvan, A., Sartori, N. Improved estimation in negative binomial regression. Statistics in Medicine, 41(13), 2403-2416 2022.
[44] Kammer, M., Dunkler, D., Michiels, S., Heinze, G. Evaluating methods for Lasso selective inference in biomedical research: A comparative simulation study. BMC Medical Research Methodology, 20, 108 2020.
[45] Taylor, J., Tibshirani, R. Post-selection inference for ℓ1-penalized likelihood models. The Canadian Journal of Statistics, 46(1), 41-61 2018.
[46] Zhang, D., Khalili, A., Asgharian, M. Post-model-selection inference in linear regression models: An integrated review. Statistics Surveys, 16, 86-136 2022.
[47] Lloyd-Smith, J. O. Maximum likelihood estimation of the negative binomial dispersion parameter for highly overdispersed data, with applications to infectious diseases. PLoS ONE, 2(2), e180 2007.
[48] Saha, K. K., Paul, S. R. Interval estimation of the negative binomial dispersion parameter. Journal of Statistical Computation and Simulation, 81(9), 1109-1122 2011.
[49] van de Geer, S., Bühlmann, P., Zhou, S. The adaptive and the thresholded lasso for potentially misspecified models (and a lower bound for the lasso). Electronic Journal of Statistics, 5, 688-749 2011.
[50] Kauermann, G., Carroll, R. J. A note on the efficiency of sandwich covariance matrix estimation. Journal of the American Statistical Association, 96(456), 1387-1396 2001.
Cite This Article
  • APA Style

    Mutunga, T., Salim, A., Kihara, P., Okenye, J. (2026). Feature Selection in Over-dispersed Binary and Count Data Models Using Penalized Optimal Estimating Functions. American Journal of Theoretical and Applied Statistics, 15(5), 276-302. https://doi.org/10.11648/j.ajtas.20261505.17

    Copy | Download

    ACS Style

    Mutunga, T.; Salim, A.; Kihara, P.; Okenye, J. Feature Selection in Over-dispersed Binary and Count Data Models Using Penalized Optimal Estimating Functions. Am. J. Theor. Appl. Stat. 2026, 15(5), 276-302. doi: 10.11648/j.ajtas.20261505.17

    Copy | Download

    AMA Style

    Mutunga T, Salim A, Kihara P, Okenye J. Feature Selection in Over-dispersed Binary and Count Data Models Using Penalized Optimal Estimating Functions. Am J Theor Appl Stat. 2026;15(5):276-302. doi: 10.11648/j.ajtas.20261505.17

    Copy | Download

  • @article{10.11648/j.ajtas.20261505.17,
      author = {Timothy Mutunga and Ali Salim and Pius Kihara and Justine Okenye},
      title = {Feature Selection in Over-dispersed Binary and Count Data Models Using Penalized Optimal Estimating Functions},
      journal = {American Journal of Theoretical and Applied Statistics},
      volume = {15},
      number = {5},
      pages = {276-302},
      doi = {10.11648/j.ajtas.20261505.17},
      url = {https://doi.org/10.11648/j.ajtas.20261505.17},
      eprint = {https://article.sciencepublishinggroup.com/pdf/10.11648.j.ajtas.20261505.17},
      abstract = {Generalized linear models (GLMs) remain a core class of supervised machine learning models for binary and count responses, with feature selection commonly carried out through penalized likelihood or quasi-likelihood methods. This paper develops a feature-selection framework based on penalized Optimal Estimating Functions (OEFs), which attain Godambe optimality within a class of unbiased estimating functions and incorporate higher-order moment information without requiring full likelihood specification. Ridge, LASSO, Adaptive LASSO and SCAD penalties are introduced at the regression estimating-equation level, while dispersion is estimated jointly through an unpenalized OEF. Hyperparameters are selected using cross-validated estimating-function loss with prespecified edge and stability rules. Monte Carlo simulations with 500 replications examine Beta-Binomial and Negative-Binomial regression under moderate and high overdispersion and sparse and moderately dense signals. The results do not show uniform superiority of penalization. The unpenalized OEF generally provides the strongest coefficient coverage and RMSE benchmark, whereas the penalized OEFs provide sparse feature selection with model-dependent trade-offs between sensitivity, specificity and interval calibration. Simple post-selection refitting reduces shrinkage bias in some settings but does not restore nominal coverage because selection uncertainty remains. Penalized OEF is therefore presented as a competitive alternative when sparse selection and joint mean-dispersion estimation are both required, rather than as a uniformly better estimator.},
     year = {2026}
    }
    

    Copy | Download

  • TY  - JOUR
    T1  - Feature Selection in Over-dispersed Binary and Count Data Models Using Penalized Optimal Estimating Functions
    AU  - Timothy Mutunga
    AU  - Ali Salim
    AU  - Pius Kihara
    AU  - Justine Okenye
    Y1  - 2026/09/24
    PY  - 2026
    N1  - https://doi.org/10.11648/j.ajtas.20261505.17
    DO  - 10.11648/j.ajtas.20261505.17
    T2  - American Journal of Theoretical and Applied Statistics
    JF  - American Journal of Theoretical and Applied Statistics
    JO  - American Journal of Theoretical and Applied Statistics
    SP  - 276
    EP  - 302
    PB  - Science Publishing Group
    SN  - 2326-9006
    UR  - https://doi.org/10.11648/j.ajtas.20261505.17
    AB  - Generalized linear models (GLMs) remain a core class of supervised machine learning models for binary and count responses, with feature selection commonly carried out through penalized likelihood or quasi-likelihood methods. This paper develops a feature-selection framework based on penalized Optimal Estimating Functions (OEFs), which attain Godambe optimality within a class of unbiased estimating functions and incorporate higher-order moment information without requiring full likelihood specification. Ridge, LASSO, Adaptive LASSO and SCAD penalties are introduced at the regression estimating-equation level, while dispersion is estimated jointly through an unpenalized OEF. Hyperparameters are selected using cross-validated estimating-function loss with prespecified edge and stability rules. Monte Carlo simulations with 500 replications examine Beta-Binomial and Negative-Binomial regression under moderate and high overdispersion and sparse and moderately dense signals. The results do not show uniform superiority of penalization. The unpenalized OEF generally provides the strongest coefficient coverage and RMSE benchmark, whereas the penalized OEFs provide sparse feature selection with model-dependent trade-offs between sensitivity, specificity and interval calibration. Simple post-selection refitting reduces shrinkage bias in some settings but does not restore nominal coverage because selection uncertainty remains. Penalized OEF is therefore presented as a competitive alternative when sparse selection and joint mean-dispersion estimation are both required, rather than as a uniformly better estimator.
    VL  - 15
    IS  - 5
    ER  - 

    Copy | Download

Author Information
  • Sections