Q1
Single choice
A machine learning (ML) specialist must develop a classification model for a financial services company. A domain expert provides the dataset, which is tabular with 10,000 rows and 1,020 features. During exploratory data analysis, the specialist finds no missing values and a small percentage of duplicate rows.
There are correlation scores of > 0.9 for 200 feature pairs. The mean value of each feature is similar to its
50th percentile.
Which feature engineering strategy should the ML specialist use with Amazon SageMaker?