
PREDICTING YOUTUBE VIDEO VIEWS USING SUPERVISED REGRESSION MACHINE LEARNING | IJCT Volume 13 – Issue 5 | IJCT-V13I5P38
IJCT
International Journal of Computer Techniques
ISSN 2394-2231 · Peer-Reviewed · Open Access
📚 Volume 13, Issue 5
📅 September 10, 2026
📄 Pages 338–344
🔖 ID: IJCT-V13I5P38
Table of Contents
TogglePREDICTING YOUTUBE VIDEO VIEWS USING SUPERVISED REGRESSION MACHINE LEARNING
Author(s)
Sakshi Bankar, Shruti Holkar, Shraddha Chavan
Abstract
Predicting the popularity of online video content is an important machine learning problem because video views are influenced by multiple measurable characteristics such as engagement, publication timing, and platform metadata. This research presents a supervised learning approach for predicting the number of views received by trending YouTube videos in India. The study uses the INvideos.csv dataset from Kaggle, which contains 37,352 records and 16 original columns. The target variable is views, a continuous numerical quantity, making the task a regression problem. During preprocessing, identifiers and free-text fields were removed, while publication timestamps were transformed into year, month, weekday, and hour features. Missing records were removed, and the data was divided into training and testing sets using an 80:20 split. Four regression algorithms—Linear Regression, Decision Tree Regression, Random Forest Regression, and K-Nearest Neighbors Regression—were trained and evaluated using Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R². Random Forest produced the best performance, achieving an R² of 0.968, compared with 0.927 for Decision Tree, 0.922 for K-Nearest Neighbors, and 0.763 for Linear Regression. The results demonstrate that ensemble-based regression can effectively estimate YouTube video views from structured metadata and engagement indicators. However, because likes, dislikes, and comments may be recorded after publication, the model should be interpreted as an estimation of popularity using available metadata rather than a strictly pre-publication forecasting system.
Keywords
YouTube analytics, machine learning, supervised learning, regression, video views, Random Forest, Kaggle, predictive modeling, engagement metrics
Conclusion
This research demonstrates a complete supervised regression workflow for predicting YouTube video views using the Kaggle Trending YouTube Video Statistics dataset for India. The 37,352-record INvideos.csv dataset was prepared by removing identifiers and free-text fields, extracting temporal features from publish_time, handling incomplete rows, separating the target variable, and scaling the predictors. Four regression algorithms were trained and evaluated using MAE, RMSE, and R². Among the tested models, Random Forest achieved the best performance with an R² of 0.968, MAE of 226,778, and RMSE of 603,833 on the held-out test set. These findings show that ensemble-based machine learning can provide strong estimates of video popularity from structured metadata and engagement indicators. At the same time, the timing of engagement variables and the use of a single train-test split should be considered when interpreting the results. Overall, the study provides a practical foundation for further work in YouTube analytics, popularity prediction, and data-driven content analysis.
References
[1] Kaggle, “Trending YouTube Video Statistics,” dataset by datasnaek, including the India file
INvideos.csv.
[2] F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning
Research, vol. 12, pp. 2825–2830, 2011.
[3] W. McKinney, Python for Data Analysis, O’Reilly Media.
[4] C. R. Harris et al., “Array programming with NumPy,” Nature, vol. 585, pp. 357–362, 2020.
[5] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning, Springer.
INvideos.csv.
[2] F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning
Research, vol. 12, pp. 2825–2830, 2011.
[3] W. McKinney, Python for Data Analysis, O’Reilly Media.
[4] C. R. Harris et al., “Array programming with NumPy,” Nature, vol. 585, pp. 357–362, 2020.
[5] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning, Springer.
📋 How to Cite This Paper
Sakshi Bankar, Shruti Holkar, Shraddha Chavan (2026). PREDICTING YOUTUBE VIDEO VIEWS USING SUPERVISED REGRESSION MACHINE LEARNING. International Journal of Computer Techniques, 13(5), 338–344. ISSN: 2394-2231. DOI: https://doi.org/10.5281/zenodo.22694731
Related Posts:









