Machine Learning: Unsupervised Learning, Neural Networks, and Reinforcement Learning
Thông tin đề
- Môn
- Machine Learning
- Kỳ thi
- University
- Số câu
- 20 câu
- Thời gian
- 40 phút
- Đáp án
- ✓ Có giải thích
- Trường
- Đại học Kinh tế – Luật (UEL)
Nội dung đề (20 câu)
- Câu 1.
In the context of Principal Components Analysis (PCA), a composite variable is best described as:
- A.
A variable that is directly measured and observed in the original dataset
- B.
A variable that combines two or more variables that are statistically strongly related to each other
- C.
A variable that has zero correlation with all other variables in the dataset
- D.
A variable that is always standardized to have a unit variance
- A.
- Câu 2.
In PCA, an eigenvalue associated with an eigenvector represents:
- A.
The number of features in the original dataset
- B.
The direction of maximum variance in the data
- C.
The proportion of total variance in the initial data that is explained by each eigenvector
- D.
The number of principal components that must be retained in the model
- A.
- Câu 3.
How does the PCA algorithm order the principal components?
- A.
Randomly, to ensure unbiased selection of components
- B.
From lowest to highest according to their eigenvalues
- C.
From highest to lowest according to their eigenvalues (i.e., in terms of usefulness in explaining total variance)
- D.
Alphabetically based on the names of the original features
- A.
- Câu 4.
What is the primary purpose of a scree plot in PCA?
- A.
To visualize the clusters formed in a dataset
- B.
To determine the optimal number of principal components to retain
- C.
To identify outliers in the original dataset
- D.
To compute the eigenvalues of the covariance matrix
- A.
- Câu 5.
Before applying PCA, the data should be standardized so that:
- A.
The mean of each series is 0 and the standard deviation is 1
- B.
All features are converted to categorical variables
- C.
The total variance of the dataset is set to 0
- D.
The covariance matrix becomes a vector of zeros
- A.
- Câu 6.
A main drawback of PCA is that:
- A.
It requires the data to be labeled in order to function
- B.
It always retains all original features, offering no dimensionality reduction
- C.
The resulting principal components typically cannot be easily labeled or directly interpreted by the analyst
- D.
It only works with datasets containing fewer than five features
- A.
- Câu 7.
A good clustering of a dataset is characterized by:
- A.
High intra-cluster distance and low inter-cluster distance
- B.
Low intra-cluster distance (cohesion) and high inter-cluster distance (separation)
- C.
Equal variance within all clusters regardless of cluster size
- D.
All clusters having the same number of observations
- A.
- Câu 8.
A commonly used distance measure in clustering, defined as the straight-line distance between two points, is:
- A.
Manhattan distance
- B.
Cosine similarity
- C.
Euclidean distance
- D.
Hamming distance
- A.
- Câu 9.
In K-means clustering, the number of clusters k is:
- A.
Determined automatically by the algorithm without any user input
- B.
A model hyperparameter that must be specified before running the algorithm
- C.
Always equal to the number of features in the dataset
- D.
Equal to the sample size of the dataset
- A.
- Câu 10.
The K-means algorithm has converged when:
- A.
All centroids coincide at the origin
- B.
The maximum number of iterations has been reached
- C.
No observation is reassigned to a new cluster (i.e., centroids no longer need to be recalculated)
- D.
The variance within each cluster is exactly zero
- A.
- Câu 11.
In K-means clustering, a centroid is best described as:
- A.
The observation that is farthest from all other observations in a cluster
- B.
The average value of the observations assigned to a cluster
- C.
The first observation included in the dataset
- D.
The mode of all categorical features within a cluster
- A.
- Câu 12.
A potential issue with the K-means algorithm is that:
- A.
The final cluster assignment can depend on the initial location of the centroids
- B.
It cannot handle datasets with more than 100 observations
- C.
It always produces clusters of equal size
- D.
It requires labeled data for training
- A.
- Câu 13.
Agglomerative hierarchical clustering is best described as a method that:
- A.
Starts with all observations in a single cluster and progressively partitions them
- B.
Begins with each observation as its own cluster and iteratively merges the two closest clusters
- C.
Requires the analyst to pre-specify the exact number of clusters before any computation
- D.
Can only be applied to datasets with exactly two features
- A.
- Câu 14.
Which of the following best characterizes the divisive approach to hierarchical clustering?
- A.
It starts with all observations in one cluster and progressively partitions them into smaller clusters
- B.
It merges the two closest clusters until only one cluster remains
- C.
It begins with k randomly chosen centroids and assigns each observation to the nearest centroid
- D.
It requires labeled training data for each observation
- A.
- Câu 15.
In a dendrogram, the height of each arch (the horizontal line connecting two vertical dendrites) represents:
- A.
The number of observations contained in the cluster
- B.
The distance (dissimilarity) between the two clusters being combined
- C.
The number of features used to construct the dendrogram
- D.
The total variance explained by the clusters
- A.
- Câu 16.
A basic neural network architecture consists of:
- A.
Only an input layer and an output layer, with no intermediate processing
- B.
An input layer, one or more hidden layers, and an output layer
- C.
Multiple input layers and a single hidden layer, with no output layer
- D.
A single layer of nodes connected in a tree structure
- A.
- Câu 17.
The activation function in a neural network node serves to:
- A.
Always produce a strictly linear output regardless of input
- B.
Transform the total net input into the final output of the node in a non-linear way
- C.
Compute the eigenvalues of the input feature matrix
- D.
Select the most important features for the model
- A.
- Câu 18.
In the training of a neural network, the process of adjusting weights by working backward through the network layers to reduce total error is called:
- A.
Forward propagation
- B.
Backward propagation
- C.
Dimensionality reduction
- D.
Cluster assignment
- A.
- Câu 19.
A deep neural network (DNN) is best described as:
- A.
A neural network with only one hidden layer
- B.
A neural network with many hidden layers (at least 2, often more than 20)
- C.
A neural network that uses only linear activation functions
- D.
A neural network that does not require any training
- A.
- Câu 20.
A key characteristic of reinforcement learning, distinguishing it from supervised learning, is that it:
- A.
Requires direct labeled data for every observation in the training set
- B.
Learns by having the agent test new actions, observe its environment, and reuse previous experiences through trials and errors
- C.
Always produces identical results regardless of the sequence of actions taken
- D.
Can only be applied to numerical prediction problems, not classification
- A.
Đáp án và giải thích từng câu có trong chế độ .