References

Adomavicius, G., & Tuzhilin, A. (2005). Toward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions. IEEE Transactions on Knowledge and Data Engineering, 17(6), 734–749.

Agrawal, R., Imieliński, T., & Swami, A. (1993). Mining association rules between sets of items in large databases. In P. Buneman & S. Jajodia (Eds.), Proceedings of the 1993 ACM SIGMOD International Conference on Management of Data (pp. 207–216). ACM. https://doi.org/10.1145/170035.170072

Agrawal, R., & Srikant, R. (1994). Fast algorithms for mining association rules in large databases. Proceedings of the 20th International Conference on Very Large Data Bases (VLDB), 487–499.

Agresti, A. (2012). Categorical data analysis (3rd ed.). Wiley.

Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27(1), 17–21. https://doi.org/10.1080/00031305.1973.10478966

Bayes, T. (1763). An essay towards solving a problem in the doctrine of chances. Philosophical Transactions of the Royal Society of London, 53, 370–418. https://doi.org/10.1098/rstl.1763.0053

Bellman, R. (1961). Adaptive control processes: A guided tour. Princeton University Press.

Blattberg, R. C., Kim, B.-D., & Neslin, S. A. (2008). Database marketing: Analyzing and managing customers (International Series in Quantitative Marketing, Vol. 18). Springer. https://doi.org/10.1007/978-0-387-72579-6

Breiman, L. (1996). Bagging predictors. Machine Learning, 24(2), 123–140. https://doi.org/10.1007/BF00058655

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). Classification and regression trees. Chapman & Hall/CRC.

Casella, G., & Berger, R. L. (2002). Statistical inference (2nd ed.). Duxbury.

Cleveland, W. S. (1979). Robust locally weighted regression and smoothing scatterplots. Journal of the American Statistical Association, 74(368), 829–836. https://doi.org/10.1080/01621459.1979.10481038

Cleveland, W. S., & Devlin, S. J. (1988). Locally weighted regression: An approach to regression analysis by local fitting. Journal of the American Statistical Association, 83(403), 596–610.

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104

Cohen, J., Cohen, P., West, S. G., & Aiken, L. S. (2003). Applied multiple regression/correlation analysis for the behavioral sciences. Routledge.

De Moivre, A. (1733). The doctrine of chances: Or, A method of calculating the probabilities of events in play (1st ed.). W. Pearson.

Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD’96) (pp. 226–231). AAAI Press.

Everitt, B., Landau, S., Leese, M., & Stahl, D. (2011). Cluster analysis (5th ed.). Wiley.

Fader, P. S., & Hardie, B. G. S. (2009). Probability models for customer-base analysis. Journal of Interactive Marketing, 23(1), 61–69. https://doi.org/10.1016/j.intmar.2008.11.003

Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning (ICML 2017) (pp. 1126–1135). PMLR. https://proceedings.mlr.press/v70/finn17a.html

Freedman, D., Pisani, R., & Purves, R. (2007). Statistics (4th ed.). W. W. Norton & Company.

Friedl, J. E. F. (2006). Mastering regular expressions (3rd ed.). O’Reilly Media.

Galton, F. (1886). Regression toward mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263. https://doi.org/10.2307/2841583

Gauss, C. F. (1809). Theoria motus corporum coelestium in sectionibus conicis solem ambientium [Theory of the motion of the heavenly bodies moving about the sun in conic sections]. Perthes et Besser.

Gelman, A., Hill, J., & Vehtari, A. (2020). Regression and other stories. Cambridge University Press.

GitHub, Inc. (2024). GitHub [Computer software]. https://github.com

Gohel, A. (2026). flextable: Functions for tabular reporting (R package version 0.7.6) [Computer software]. https://cran.r-project.org/package=flextable

Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

Greene, W. H. (2018). Econometric analysis (8th ed.). Pearson.

Han, J., Kamber, M., & Pei, J. (2012). Data mining: Concepts and techniques (3rd ed.). Morgan Kaufmann.

Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7

Hunter, J. D. (2007). Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9(3), 90–95. https://doi.org/10.1109/MCSE.2007.55

Hyndman, R. J., & Athanasopoulos, G. (2021). Forecasting: Principles and practice (3rd ed.). OTexts. https://otexts.com/fpp3/

Irizarry, R. A. (2024). Introduction to data science: Data wrangling and visualization with R (2nd ed.). Chapman & Hall/CRC.

Jain, A. K. (2010). Data clustering: 50 years beyond k-means. Pattern Recognition Letters, 31(8), 651–666.

Johnson, S. C. (1967). Hierarchical clustering schemes. Psychometrika, 32(3), 241–254. https://doi.org/10.1007/BF02289588

Jolliffe, I. T. (2002). Principal component analysis. Springer.

Jupyter Development Team. (n.d.). Jupyter. Retrieved May 12, 2026, from https://jupyter.org/

Kaufman, L., & Rousseeuw, P. J. (2009). Finding groups in data: An introduction to cluster analysis (2nd ed.). Wiley.

Ketchen, D. J., Jr., & Shook, C. L. (1996). The application of cluster analysis in strategic management research: An analysis and critique. Strategic Management Journal, 17(6), 441–458.

Kuhn, M., & Johnson, K. (2013). Applied predictive modeling. Springer.

Kuhn, M., & Johnson, K. (2019). Feature engineering and selection: A practical approach for predictive models. CRC Press.

Kuhn, M., & Silge, J. (2022). Tidymodels: A framework for modeling and machine learning using tidyverse principles. O’Reilly Media.

Lantz, B. (2023). Machine learning with R: Learn techniques for building and improving machine learning models, from data preparation to model tuning, evaluation, and working with big data (4th ed.). Packt Publishing Limited. https://www.packtpub.com/en-gb/product/machine-learning-with-r-9781801076050

LeCun, Y., & He, K. (2022). Deep learning. Nature, 604(7900), 921–930. https://doi.org/10.1038/s41586-022-04455-8

Legendre, A.-M. (1805). Nouvelles méthodes pour la détermination des orbites des comètes [New methods for the determination of cometary orbits]. F. Didot.

Lloyd, S. P. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129–137.

MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In L. M. LeCam & J. Neyman (Eds.), Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability (Vol. 1, pp. 281–297). University of California Press.

McCallum, A., & Nigam, K. (1998). A comparison of event models for Naive Bayes text classification. In AAAI-98 Workshop on Learning for Text Categorization (Vol. 752, pp. 41–48). AAAI Press.

McElreath, R. (2020). Statistical rethinking: A Bayesian course with examples in R and Stan (2nd ed.). CRC Press.

Microsoft. (n.d.). Visual Studio Code. https://code.visualstudio.com/

Mitchell, T. M. (1997). Machine learning. McGraw-Hill.

Montgomery, D. C., Peck, E. A., & Vining, G. G. (2021). Introduction to linear regression analysis (6th ed.). Wiley.

Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). W. H. Freeman & Company.

Moro, S., Rita, P., & Laureano, R. (2018). Hotel booking demand datasets. Data in Brief, 22, 41–49.

Murtagh, F., & Contreras, P. (2012). Algorithms for hierarchical clustering: An overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2(1), 86–97. https://doi.org/10.1002/widm.53

NumPy Developers. (n.d.). NumPy (Version 2.x). https://numpy.org/

Ockham, W. of. (1990). Philosophical writings: A selection (P. Boehner, Ed.). Hackett Publishing Company. (Original work published 14th century)

pandas Development Team. (n.d.). pandas documentation. https://pandas.pydata.org/docs/

Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.

Pearson, K. (1896). Mathematical contributions to the theory of evolution. III. Regression, heredity, and panmixia. Philosophical Transactions of the Royal Society of London. Series A, 187, 253–318. https://doi.org/10.1098/rsta.1896.0007

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Python Software Foundation. (n.d.). datetime — Basic date and time types. Python Documentation. https://docs.python.org/3/library/datetime.html

Python Software Foundation. (n.d.). History of Python. https://www.python.org/doc/essays/blurb/

Python Software Foundation. (n.d.). math — Mathematical functions. Python Standard Library. Retrieved May 12, 2026, from https://docs.python.org/3/library/math.html

Python Software Foundation. (n.d.). Python documentation. https://docs.python.org/

R Core Team. (2026). mtcars: Motor Trend car road tests (1974) [Data set]. In R: A language and environment for statistical computing (Version 5.6.0). R Foundation for Statistical Computing. https://www.r-project.org/

Schafer, J. B., Konstan, J. A., & Riedl, J. (2001). E-commerce recommendation applications. Data Mining and Knowledge Discovery, 5(1–2), 115–153.

Seabold, S., & Perktold, J. (2010). Statsmodels: Econometric and statistical modeling with Python. Proceedings of the 9th Python in Science Conference, 92–96. https://doi.org/10.25080/Majora-92bf1922-011

Sievert, C. (2023). plotly: Create interactive web graphics via ‘plotly.js’ (R package version 4.11.1) [Computer software]. CRAN. https://CRAN.R-project.org/package=plotly

Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

Swets, J. A. (1988). Measuring the accuracy of diagnostic systems. Science, 240(4857), 1285–1293. https://doi.org/10.1126/science.3287615

Tan, P.-N., Steinbach, M., & Kumar, V. (2019). Introduction to data mining (2nd ed.). Pearson.

Taunk, K., De, S., & Verma, S. (2019). A brief review of nearest neighbor algorithm for learning and classification. Proceedings of the 2019 International Conference on Intelligent Computing and Control Systems (ICCS). https://doi.org/10.1109/ICCS45141.2019.9065747

Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley.

Urdan, T. C. (2022). Statistics in plain English (5th ed.). Routledge. https://doi.org/10.4324/9781003196582

Vaughan, D. (2017). Statistics for the pharmaceutical sciences. CRC Press.

Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., … SciPy 1.0 Contributors. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261–272. https://doi.org/10.1038/s41592-019-0686-2

Waskom, M. (2021). seaborn: Statistical data visualization (Version 0.11.2) [Computer software]. https://doi.org/10.5281/zenodo.4569847

Wedel, M., & Kamakura, W. A. (2000). Market segmentation: Conceptual and methodological foundations (2nd ed.). Springer.

Wickham, H. (2014). Tidy data. Journal of Statistical Software, 59(10), 1–23. https://doi.org/10.18637/jss.v059.i10

Wolpert, D. H. (1996). The lack of a priori distinctions between learning algorithms. Neural Computation, 8(7), 1341–1390. https://doi.org/10.1162/neco.1996.8.7.1341

Zhang, H. (2004). The optimality of Naive Bayes. In Proceedings of the Seventeenth International Florida Artificial Intelligence Research Society Conference (FLAIRS 2004) (pp. 562–567). AAAI Press.

Zhu, X., & Goldberg, A. B. (2009). Introduction to semi-supervised learning. Synthesis Lectures on Artificial Intelligence and Machine Learning, 3(1), 1–130. https://doi.org/10.2200/S00196ED1V01Y200906AIM006