Data Analytics and Data Mining – 100 Tough MCQs
1. Which statement best distinguishes data mining from traditional query processing?
A. Query processing discovers hidden patterns, while data mining retrieves known facts.
B. Data mining retrieves predefined reports, while query processing predicts future events.
C. Data mining discovers previously unknown patterns, while query processing retrieves explicitly stored information.
D. Both perform identical operations but on different databases.
Answer: C
2. Which of the following tasks is NOT considered a predictive data mining task?
A. Classification
B. Regression
C. Clustering
D. Forecasting
Answer: C
3. In the CRISP-DM methodology, the phase immediately after Business Understanding is:
A. Deployment
B. Data Preparation
C. Data Understanding
D. Modeling
Answer: C
4. Which property of a data warehouse ensures that historical data is retained and not updated?
A. Subject-oriented
B. Time-variant
C. Integrated
D. Non-volatile
Answer: D
5. The primary objective of dimensionality reduction is to:
A. Increase the number of features
B. Improve storage only
C. Reduce redundancy while preserving useful information
D. Remove categorical variables
Answer: C
6. Which similarity measure is most suitable for market basket analysis involving binary attributes?
A. Euclidean Distance
B. Manhattan Distance
C. Cosine Similarity
D. Jaccard Coefficient
Answer: D
7. In association rule mining, confidence measures:
A. Frequency of occurrence of an itemset
B. Probability that the consequent occurs given the antecedent
C. Independence between variables
D. Number of transactions
Answer: B
8. Which algorithm is specifically designed for mining frequent itemsets without candidate generation?
A. Apriori
B. FP-Growth
C. ID3
D. KNN
Answer: B
9. Which impurity measure is used by the CART algorithm?
A. Entropy
B. Information Gain
C. Gini Index
D. Chi-square
Answer: C
10. Information Gain tends to favor:
A. Binary attributes
B. Continuous variables
C. Attributes having many distinct values
D. Balanced classes
Answer: C
11. Which validation method leaves exactly one observation out for testing during each iteration?
A. Holdout Validation
B. Stratified Sampling
C. Leave-One-Out Cross Validation
D. Bootstrap
Answer: C
12. Which metric is most appropriate when false negatives are significantly more costly than false positives?
A. Accuracy
B. Precision
C. Recall
D. Specificity
Answer: C
13. The "curse of dimensionality" mainly affects:
A. SQL queries
B. Distance-based algorithms
C. OLAP cubes
D. ETL pipelines
Answer: B
14. Which clustering algorithm does NOT require specifying the number of clusters beforehand?
A. K-Means
B. DBSCAN
C. K-Medoids
D. Bisecting K-Means
Answer: B
15. Which algorithm is particularly sensitive to the initial selection of centroids?
A. Decision Tree
B. Naïve Bayes
C. K-Means
D. Random Forest
Answer: C
16. In Principal Component Analysis (PCA), principal components are:
A. Always correlated
B. Linear combinations of original variables
C. Original variables after normalization
D. Decision boundaries
Answer: B
17. Which preprocessing technique is most suitable for handling highly skewed numerical data?
A. One-Hot Encoding
B. Log Transformation
C. Label Encoding
D. Dummy Coding
Answer: B
18. Which of the following is an unsupervised learning technique?
A. Logistic Regression
B. Support Vector Machine
C. Hierarchical Clustering
D. Random Forest
Answer: C
19. Which statement about Naïve Bayes is TRUE?
A. It assumes attributes are mutually exclusive.
B. It assumes conditional independence among predictors.
C. It requires numerical attributes only.
D. It cannot classify multiclass problems.
Answer: B
20. Which kernel allows Support Vector Machines to classify non-linearly separable data?
A. Identity Kernel
B. Polynomial or RBF Kernel
C. Mean Kernel
D. Binary Kernel
Answer: B
21. Which ensemble learning technique primarily aims to reduce variance?
A. Boosting
B. Bagging
C. Gradient Descent
D. PCA
Answer: B
22. In Random Forest, randomness is introduced by:
A. Using different distance metrics
B. Selecting random subsets of data and features
C. Changing class labels
D. Randomly removing decision trees
Answer: B
23. Which data quality issue occurs when the same real-world entity appears multiple times with different identifiers?
A. Missing Values
B. Duplicate Records
C. Inconsistent Encoding
D. Data Drift
Answer: B
24. Which OLAP operation transforms detailed data into summarized data?
A. Drill-down
B. Slice
C. Roll-up
D. Pivot
Answer: C
25. Which statement about Hadoop MapReduce is correct?
A. Mapping aggregates final results.
B. Reduce executes before Map.
C. Mapping processes input into intermediate key-value pairs before reduction.
D. MapReduce requires a relational database.
Answer: C
26. Which of the following is not one of the "3 Vs" originally associated with Big Data?
A. Volume
B. Velocity
C. Variety
D. Validity
Answer: D
27. In Apache Spark, which component is primarily responsible for in-memory distributed computation?
A. HDFS
B. RDD (Resilient Distributed Dataset)
C. Hive
D. Sqoop
Answer: B
28. Which ETL operation is responsible for resolving inconsistent date formats before loading data into a warehouse?
A. Extraction
B. Transformation
C. Loading
D. Replication
Answer: B
29. Which normalization method scales data into the interval [0,1]?
A. Z-score Normalization
B. Decimal Scaling
C. Min-Max Normalization
D. Mean Centering
Answer: C
30. A model achieves 99% accuracy on a dataset where 99% of observations belong to one class. This is an example of:
A. High bias
B. Data leakage
C. Misleading accuracy due to class imbalance
D. Underfitting
Answer: C
31. Which metric combines Precision and Recall into a single harmonic mean?
A. ROC Score
B. Specificity
C. F1 Score
D. Lift
Answer: C
32. ROC curves plot:
A. Precision against Recall
B. Recall against Accuracy
C. True Positive Rate against False Positive Rate
D. Sensitivity against Specificity
Answer: C
33. An AUC (Area Under Curve) value of 0.5 indicates:
A. Perfect classifier
B. Random guessing
C. Excellent classifier
D. Severe overfitting
Answer: B
34. Which statistical measure is least affected by extreme outliers?
A. Mean
B. Variance
C. Median
D. Standard Deviation
Answer: C
35. Which technique is commonly used to detect multicollinearity among predictor variables?
A. Silhouette Score
B. Variance Inflation Factor (VIF)
C. Lift Ratio
D. Entropy
Answer: B
36. Which clustering evaluation metric does not require ground-truth labels?
A. Accuracy
B. Silhouette Coefficient
C. Precision
D. Recall
Answer: B
37. In DBSCAN, points that are neither core points nor border points are called:
A. Seed Points
B. Noise Points
C. Pivot Points
D. Anchor Points
Answer: B
38. Which of the following algorithms is deterministic (produces the same result every run with identical data)?
A. Standard K-Means
B. Random Forest
C. Hierarchical Agglomerative Clustering
D. Bagging
Answer: C
39. Which distance measure is most suitable when variables have different units without prior normalization?
A. Euclidean Distance
B. Manhattan Distance
C. Mahalanobis Distance
D. Hamming Distance
Answer: C
40. Which technique reduces overfitting by penalizing large regression coefficients?
A. Bootstrap Sampling
B. Regularization
C. One-Hot Encoding
D. Cross Join
Answer: B
41. LASSO regression differs from Ridge regression because it:
A. Uses no penalty term
B. Can shrink some coefficients exactly to zero
C. Always produces larger coefficients
D. Works only for classification
Answer: B
42. Which assumption is essential for Linear Regression but not for Decision Trees?
A. Presence of categorical variables
B. Linear relationship between predictors and response
C. Missing values
D. Binary output
Answer: B
43. In time series analysis, autocorrelation refers to:
A. Correlation between independent variables
B. Correlation of a variable with its own past values
C. Correlation between training and testing data
D. Correlation between clusters
Answer: B
44. Which forecasting model explicitly includes trend, seasonality, and error components?
A. K-Means
B. Holt-Winters Exponential Smoothing
C. Naïve Bayes
D. Apriori
Answer: B
45. Which SQL operation is most commonly used during ETL to combine records from multiple tables?
A. DROP
B. ALTER
C. JOIN
D. TRUNCATE
Answer: C
46. Which NoSQL database type is best suited for storing graph relationships such as social networks?
A. Key-Value Store
B. Column Family Database
C. Graph Database
D. Document Store
Answer: C
47. Which Apache Hadoop ecosystem component provides SQL-like querying over distributed data?
A. Pig
B. Hive
C. Flume
D. Oozie
Answer: B
48. Which of the following is an example of descriptive analytics?
A. Predicting customer churn
B. Estimating next year's revenue
C. Summarizing last quarter's sales by region
D. Optimizing delivery routes
Answer: C
49. Which type of analytics answers the question "What should we do?"
A. Descriptive Analytics
B. Diagnostic Analytics
C. Predictive Analytics
D. Prescriptive Analytics
Answer: D
50. Which statement about data governance is most accurate?
A. It deals only with database security.
B. It focuses only on backup and recovery.
C. It establishes policies, standards, ownership, and accountability for data assets.
D. It is synonymous with data mining.
Answer: C
51. In association rule mining, a Lift value greater than 1 indicates:
A. The antecedent and consequent are negatively correlated.
B. The antecedent and consequent occur independently.
C. The occurrence of the antecedent increases the likelihood of the consequent.
D. The rule has zero confidence.
Answer: C
52. Which association rule measure is symmetric, unlike confidence?
A. Support
B. Lift
C. Confidence
D. Conviction
Answer: B
53. Which algorithm is specifically designed for anomaly detection by randomly partitioning data?
A. K-Means
B. Isolation Forest
C. Apriori
D. FP-Growth
Answer: B
54. Local Outlier Factor (LOF) identifies anomalies based on:
A. Global mean deviation
B. Density relative to neighboring observations
C. Euclidean distance from origin
D. Cluster centroids only
Answer: B
55. Concept drift refers to:
A. Missing values in a dataset.
B. Gradual or sudden changes in the statistical properties of data over time.
C. Loss of database indexes.
D. Duplicate records in a warehouse.
Answer: B
56. Which technique is commonly used to handle highly imbalanced classification datasets?
A. PCA
B. SMOTE
C. K-Means
D. Min-Max Scaling
Answer: B
57. SMOTE generates:
A. Random noise variables.
B. Synthetic minority class samples.
C. Additional majority class records.
D. Duplicate observations only.
Answer: B
58. Which ensemble algorithm builds trees sequentially, where each new tree attempts to correct previous errors?
A. Random Forest
B. Bagging
C. Gradient Boosting
D. Extra Trees
Answer: C
59. XGBoost improves Gradient Boosting mainly through:
A. Removing regularization.
B. Parallelization and regularization techniques.
C. Eliminating decision trees.
D. Using only categorical variables.
Answer: B
60. Which impurity measure can never be negative?
A. Entropy
B. Gini Index
C. Variance
D. All of the above
Answer: D
61. In information theory, entropy becomes zero when:
A. All classes occur with equal probability.
B. Every class has exactly two observations.
C. The dataset contains only one class.
D. The dataset has missing values.
Answer: C
62. Which theorem forms the mathematical foundation of the Naïve Bayes classifier?
A. Central Limit Theorem
B. Bayes' Theorem
C. Markov Theorem
D. Chebyshev's Theorem
Answer: B
63. Which NLP technique converts text into numerical vectors based on word frequencies?
A. PCA
B. TF-IDF
C. One-Hot Encoding of Classes
D. Z-score Scaling
Answer: B
64. In TF-IDF, the IDF component gives higher weight to words that:
A. Appear in every document.
B. Are rare across documents.
C. Are longest in length.
D. Contain numerical values.
Answer: B
65. Which web mining category focuses on analyzing hyperlink structures?
A. Web Usage Mining
B. Web Structure Mining
C. Web Content Mining
D. Stream Mining
Answer: B
66. Clickstream analysis is primarily associated with:
A. Web Usage Mining
B. Data Warehousing
C. OLTP Systems
D. Image Mining
Answer: A
67. Stream mining algorithms must primarily address:
A. Unlimited storage availability.
B. Infinite processing time.
C. Continuous arrival of data with limited memory.
D. Static datasets only.
Answer: C
68. Which property best distinguishes a Data Lake from a traditional Data Warehouse?
A. Data Lake stores only structured data.
B. Data Warehouse stores only unstructured data.
C. Data Lake can store structured, semi-structured, and unstructured data in raw form.
D. Data Warehouse requires Hadoop.
Answer: C
69. Which privacy model extends k-anonymity by ensuring diversity in sensitive attributes?
A. PCA
B. l-diversity
C. FP-Growth
D. ROC
Answer: B
70. Differential Privacy primarily aims to:
A. Compress datasets.
B. Protect individual information by adding controlled statistical noise.
C. Increase classification accuracy.
D. Improve clustering quality.
Answer: B
71. Which sampling method gives every population member an equal probability of selection?
A. Stratified Sampling
B. Cluster Sampling
C. Simple Random Sampling
D. Systematic Sampling
Answer: C
72. Which statistical test is generally used to compare the means of two independent samples?
A. Chi-Square Test
B. t-Test
C. ANOVA
D. Mann-Whitney Test
Answer: B
73. Which hypothesis error occurs when a true null hypothesis is rejected?
A. Type II Error
B. Sampling Error
C. Type I Error
D. Standard Error
Answer: C
74. A p-value less than the chosen significance level (α) indicates:
A. Accept the null hypothesis.
B. Reject the null hypothesis.
C. Increase sample size immediately.
D. The experiment is invalid.
Answer: B
75. Which statement about Explainable AI (XAI) is most accurate?
A. It aims to replace machine learning models with SQL queries.
B. It focuses on making AI model decisions understandable to humans.
C. It eliminates the need for model validation.
D. It guarantees 100% prediction accuracy.
Answer: B
76. Which OLAP operation creates a smaller cube by selecting a single value for one of the dimensions?
A. Roll-up
B. Drill-down
C. Slice
D. Pivot
Answer: C
77. Which OLAP operation changes the dimensional orientation of a data cube to provide an alternative presentation of data?
A. Drill-down
B. Pivot (Rotate)
C. Roll-up
D. Slice
Answer: B
78. Which data cube computation strategy computes only the cuboids that are actually required, thereby reducing storage and computation?
A. Full Materialization
B. No Materialization
C. Partial Materialization
D. Complete Aggregation
Answer: C
79. Which optimization technique significantly improves the efficiency of the Apriori algorithm?
A. Candidate pruning using the downward closure property
B. Increasing the minimum support threshold after every iteration
C. Sorting transactions alphabetically
D. Using Euclidean distance
Answer: A
80. Sequential Pattern Mining differs from Association Rule Mining because it considers:
A. Only numerical attributes
B. Temporal or ordered relationships among events
C. Only binary data
D. Data compression techniques
Answer: B
81. Which feature selection method evaluates variables independently of any machine learning algorithm?
A. Wrapper Method
B. Embedded Method
C. Filter Method
D. Ensemble Method
Answer: C
82. In Principal Component Analysis (PCA), the first principal component is the one that:
A. Has the smallest variance
B. Is most correlated with the target variable
C. Explains the maximum variance in the data
D. Contains only categorical variables
Answer: C
83. Which matrix is decomposed into eigenvalues and eigenvectors during Principal Component Analysis?
A. Identity Matrix
B. Covariance Matrix (or Correlation Matrix after standardization)
C. Confusion Matrix
D. Transition Matrix
Answer: B
84. Covariance differs from correlation because covariance:
A. Is always between –1 and +1
B. Is independent of the units of measurement
C. Depends on the scale (units) of the variables
D. Can never be negative
Answer: C
85. Which Apache Spark component is responsible for coordinating the execution of an application?
A. Worker Node
B. Driver Program
C. DataNode
D. NameNode
Answer: B
86. In the Hadoop ecosystem, YARN is primarily responsible for:
A. Data compression
B. Resource management and job scheduling
C. SQL querying
D. Data visualization
Answer: B
87. According to the CAP Theorem, a distributed database can guarantee at most which two of the following three properties simultaneously?
A. Capacity, Availability, Performance
B. Consistency, Availability, Partition Tolerance
C. Consistency, Accuracy, Partition Tolerance
D. Availability, Performance, Security
Answer: B
88. Which property is associated with NoSQL databases under the BASE model rather than the ACID model?
A. Atomicity
B. Immediate Consistency
C. Eventual Consistency
D. Isolation
Answer: C
89. Which SQL window function assigns the same rank to tied rows while leaving gaps in subsequent ranks?
A. ROW_NUMBER()
B. DENSE_RANK()
C. RANK()
D. NTILE()
Answer: C
90. Data lineage refers to:
A. The genealogy of employees in an organization
B. The lifecycle of data, including its origin, transformations, and movement
C. The structure of a database schema only
D. Backup history of a database
Answer: B
91. Metadata is best described as:
A. Duplicate data
B. Data about data
C. Temporary data
D. Backup data
Answer: B
92. Which phase of MLOps focuses on continuously tracking model performance after deployment?
A. Feature Engineering
B. Model Monitoring
C. Data Cleaning
D. Data Integration
Answer: B
93. Model drift occurs when:
A. The programming language changes
B. The statistical relationship between input data and target variable changes over time
C. Storage capacity decreases
D. The model file becomes corrupted
Answer: B
94. Which statement best describes the Bias–Variance Tradeoff?
A. Reducing bias always reduces variance.
B. Increasing model complexity generally reduces bias but may increase variance.
C. Bias and variance are unrelated.
D. High variance always improves generalization.
Answer: B
95. The primary objective of an A/B test is to:
A. Compress large datasets
B. Compare two alternatives using statistical evidence
C. Remove missing values
D. Build classification trees
Answer: B
96. A Bayesian Network is best described as:
A. A relational database model
B. A probabilistic graphical model representing conditional dependencies among variables
C. A clustering algorithm
D. A distributed file system
Answer: B
97. Which probability theorem is commonly used to compute the probability of an event by partitioning the sample space into mutually exclusive cases?
A. Bayes' Theorem
B. Law of Total Probability
C. Chebyshev's Inequality
D. Markov Property
Answer: B
98. Which measure is most appropriate for evaluating regression models?
A. Accuracy
B. F1-Score
C. Mean Squared Error (MSE)
D. Precision
Answer: C
99. Which statement about feature engineering is correct?
A. It is useful only for deep learning models.
B. Creating meaningful features from raw data can significantly improve model performance.
C. It always increases the number of features.
D. It is performed only after model deployment.
Answer: B
100. Which of the following best summarizes the goal of data mining?
A. To store data efficiently in databases.
B. To retrieve predefined reports from databases.
C. To discover valid, novel, useful, and understandable patterns from large datasets for decision-making.
D. To replace statistical analysis entirely.
Answer: C
1 $type={blogger}:
AKGVG is the leading audit firm in India. We are a team of proficient and dedicated chartered accountants based in New Delhi as well as other major cities in India. The Audit Services are tailored to your specific needs, and combined as relevant.
audit firms in Delhi
Post a Comment