Machine Learning: Unsupervised Learning, Neural Networks, and Reinforcement Learning

20 câu
40 phút
Có đáp án

Thông tin đề

Môn
Machine Learning
Kỳ thi
University
Số câu
20 câu
Thời gian
40 phút
Đáp án
✓ Có giải thích

Nội dung đề (20 câu)

  1. Câu 1.

    In the context of Principal Components Analysis (PCA), a composite variable is best described as:

    • A.

      A variable that is directly measured and observed in the original dataset

    • B.

      A variable that combines two or more variables that are statistically strongly related to each other

    • C.

      A variable that has zero correlation with all other variables in the dataset

    • D.

      A variable that is always standardized to have a unit variance

  2. Câu 2.

    In PCA, an eigenvalue associated with an eigenvector represents:

    • A.

      The number of features in the original dataset

    • B.

      The direction of maximum variance in the data

    • C.

      The proportion of total variance in the initial data that is explained by each eigenvector

    • D.

      The number of principal components that must be retained in the model

  3. Câu 3.

    How does the PCA algorithm order the principal components?

    • A.

      Randomly, to ensure unbiased selection of components

    • B.

      From lowest to highest according to their eigenvalues

    • C.

      From highest to lowest according to their eigenvalues (i.e., in terms of usefulness in explaining total variance)

    • D.

      Alphabetically based on the names of the original features

  4. Câu 4.

    What is the primary purpose of a scree plot in PCA?

    • A.

      To visualize the clusters formed in a dataset

    • B.

      To determine the optimal number of principal components to retain

    • C.

      To identify outliers in the original dataset

    • D.

      To compute the eigenvalues of the covariance matrix

  5. Câu 5.

    Before applying PCA, the data should be standardized so that:

    • A.

      The mean of each series is 0 and the standard deviation is 1

    • B.

      All features are converted to categorical variables

    • C.

      The total variance of the dataset is set to 0

    • D.

      The covariance matrix becomes a vector of zeros

  6. Câu 6.

    A main drawback of PCA is that:

    • A.

      It requires the data to be labeled in order to function

    • B.

      It always retains all original features, offering no dimensionality reduction

    • C.

      The resulting principal components typically cannot be easily labeled or directly interpreted by the analyst

    • D.

      It only works with datasets containing fewer than five features

  7. Câu 7.

    A good clustering of a dataset is characterized by:

    • A.

      High intra-cluster distance and low inter-cluster distance

    • B.

      Low intra-cluster distance (cohesion) and high inter-cluster distance (separation)

    • C.

      Equal variance within all clusters regardless of cluster size

    • D.

      All clusters having the same number of observations

  8. Câu 8.

    A commonly used distance measure in clustering, defined as the straight-line distance between two points, is:

    • A.

      Manhattan distance

    • B.

      Cosine similarity

    • C.

      Euclidean distance

    • D.

      Hamming distance

  9. Câu 9.

    In K-means clustering, the number of clusters k is:

    • A.

      Determined automatically by the algorithm without any user input

    • B.

      A model hyperparameter that must be specified before running the algorithm

    • C.

      Always equal to the number of features in the dataset

    • D.

      Equal to the sample size of the dataset

  10. Câu 10.

    The K-means algorithm has converged when:

    • A.

      All centroids coincide at the origin

    • B.

      The maximum number of iterations has been reached

    • C.

      No observation is reassigned to a new cluster (i.e., centroids no longer need to be recalculated)

    • D.

      The variance within each cluster is exactly zero

  11. Câu 11.

    In K-means clustering, a centroid is best described as:

    • A.

      The observation that is farthest from all other observations in a cluster

    • B.

      The average value of the observations assigned to a cluster

    • C.

      The first observation included in the dataset

    • D.

      The mode of all categorical features within a cluster

  12. Câu 12.

    A potential issue with the K-means algorithm is that:

    • A.

      The final cluster assignment can depend on the initial location of the centroids

    • B.

      It cannot handle datasets with more than 100 observations

    • C.

      It always produces clusters of equal size

    • D.

      It requires labeled data for training

  13. Câu 13.

    Agglomerative hierarchical clustering is best described as a method that:

    • A.

      Starts with all observations in a single cluster and progressively partitions them

    • B.

      Begins with each observation as its own cluster and iteratively merges the two closest clusters

    • C.

      Requires the analyst to pre-specify the exact number of clusters before any computation

    • D.

      Can only be applied to datasets with exactly two features

  14. Câu 14.

    Which of the following best characterizes the divisive approach to hierarchical clustering?

    • A.

      It starts with all observations in one cluster and progressively partitions them into smaller clusters

    • B.

      It merges the two closest clusters until only one cluster remains

    • C.

      It begins with k randomly chosen centroids and assigns each observation to the nearest centroid

    • D.

      It requires labeled training data for each observation

  15. Câu 15.

    In a dendrogram, the height of each arch (the horizontal line connecting two vertical dendrites) represents:

    • A.

      The number of observations contained in the cluster

    • B.

      The distance (dissimilarity) between the two clusters being combined

    • C.

      The number of features used to construct the dendrogram

    • D.

      The total variance explained by the clusters

  16. Câu 16.

    A basic neural network architecture consists of:

    • A.

      Only an input layer and an output layer, with no intermediate processing

    • B.

      An input layer, one or more hidden layers, and an output layer

    • C.

      Multiple input layers and a single hidden layer, with no output layer

    • D.

      A single layer of nodes connected in a tree structure

  17. Câu 17.

    The activation function in a neural network node serves to:

    • A.

      Always produce a strictly linear output regardless of input

    • B.

      Transform the total net input into the final output of the node in a non-linear way

    • C.

      Compute the eigenvalues of the input feature matrix

    • D.

      Select the most important features for the model

  18. Câu 18.

    In the training of a neural network, the process of adjusting weights by working backward through the network layers to reduce total error is called:

    • A.

      Forward propagation

    • B.

      Backward propagation

    • C.

      Dimensionality reduction

    • D.

      Cluster assignment

  19. Câu 19.

    A deep neural network (DNN) is best described as:

    • A.

      A neural network with only one hidden layer

    • B.

      A neural network with many hidden layers (at least 2, often more than 20)

    • C.

      A neural network that uses only linear activation functions

    • D.

      A neural network that does not require any training

  20. Câu 20.

    A key characteristic of reinforcement learning, distinguishing it from supervised learning, is that it:

    • A.

      Requires direct labeled data for every observation in the training set

    • B.

      Learns by having the agent test new actions, observe its environment, and reuse previous experiences through trials and errors

    • C.

      Always produces identical results regardless of the sequence of actions taken

    • D.

      Can only be applied to numerical prediction problems, not classification

Đáp án và giải thích từng câu có trong chế độ .