Recently Published
PrakKomstat_Pert2
Rahmania Nur Hafidzah (2507016011)
Clase Prueba 21 de septiembre de 2026
Este documento solo es prueba
PrakKomstat_Pert2
Rahmania Nur Hafidzah (2507016011)
TM.Komstat_Niska Pradila
Niska Pradila
2507016073
Kelompok 4B
Imputacion retail
This stage continues the data-quality and deterministic-cleaning process previously performed in Python. During that stage, missing values in Item and Price Per Unit were recovered using validated relationships between the variables.
Three variables still contain missing values: Quantity, Total Spent, and Discount Applied. However, they do not represent three independent imputation problems.
Quantity requires statistical imputation because its missing values cannot be recovered deterministically from the available information. Once Quantity has been estimated, the corresponding missing values of Total Spent can be reconstructed exactly through the validated business relationship:
Total Spent=Price Per Unit×Quantity
Therefore, Total Spent will not be statistically imputed.
Discount Applied constitutes a separate binary missing-data problem and will be evaluated independently.
The methodological strategy is therefore:
statistically impute Quantity;
reconstruct Total Spent deterministically;
statistically impute Discount Applied;
preserve all originally observed values;
validate the resulting distributions and business rules.
For Quantity, Predictive Mean Matching (PMM), CART, and Random Forest will be compared. For Discount Applied, logistic regression, CART, and Random Forest will be evaluated.
final project R Data Vis
waaahhh