11 June 2022

Data Analytics and Data Mining – 100 Tough MCQs

Posted by with 1 comment

 


Data Analytics and Data Mining – 100 Tough MCQs


1. Which statement best distinguishes data mining from traditional query processing?

A. Query processing discovers hidden patterns, while data mining retrieves known facts.
B. Data mining retrieves predefined reports, while query processing predicts future events.
C. Data mining discovers previously unknown patterns, while query processing retrieves explicitly stored information.
D. Both perform identical operations but on different databases.

Answer: C


2. Which of the following tasks is NOT considered a predictive data mining task?

A. Classification
B. Regression
C. Clustering
D. Forecasting

Answer: C


3. In the CRISP-DM methodology, the phase immediately after Business Understanding is:

A. Deployment
B. Data Preparation
C. Data Understanding
D. Modeling

Answer: C


4. Which property of a data warehouse ensures that historical data is retained and not updated?

A. Subject-oriented
B. Time-variant
C. Integrated
D. Non-volatile

Answer: D


5. The primary objective of dimensionality reduction is to:

A. Increase the number of features
B. Improve storage only
C. Reduce redundancy while preserving useful information
D. Remove categorical variables

Answer: C


6. Which similarity measure is most suitable for market basket analysis involving binary attributes?

A. Euclidean Distance
B. Manhattan Distance
C. Cosine Similarity
D. Jaccard Coefficient

Answer: D


7. In association rule mining, confidence measures:

A. Frequency of occurrence of an itemset
B. Probability that the consequent occurs given the antecedent
C. Independence between variables
D. Number of transactions

Answer: B


8. Which algorithm is specifically designed for mining frequent itemsets without candidate generation?

A. Apriori
B. FP-Growth
C. ID3
D. KNN

Answer: B


9. Which impurity measure is used by the CART algorithm?

A. Entropy
B. Information Gain
C. Gini Index
D. Chi-square

Answer: C


10. Information Gain tends to favor:

A. Binary attributes
B. Continuous variables
C. Attributes having many distinct values
D. Balanced classes

Answer: C


11. Which validation method leaves exactly one observation out for testing during each iteration?

A. Holdout Validation
B. Stratified Sampling
C. Leave-One-Out Cross Validation
D. Bootstrap

Answer: C


12. Which metric is most appropriate when false negatives are significantly more costly than false positives?

A. Accuracy
B. Precision
C. Recall
D. Specificity

Answer: C


13. The "curse of dimensionality" mainly affects:

A. SQL queries
B. Distance-based algorithms
C. OLAP cubes
D. ETL pipelines

Answer: B


14. Which clustering algorithm does NOT require specifying the number of clusters beforehand?

A. K-Means
B. DBSCAN
C. K-Medoids
D. Bisecting K-Means

Answer: B


15. Which algorithm is particularly sensitive to the initial selection of centroids?

A. Decision Tree
B. Naïve Bayes
C. K-Means
D. Random Forest

Answer: C


16. In Principal Component Analysis (PCA), principal components are:

A. Always correlated
B. Linear combinations of original variables
C. Original variables after normalization
D. Decision boundaries

Answer: B


17. Which preprocessing technique is most suitable for handling highly skewed numerical data?

A. One-Hot Encoding
B. Log Transformation
C. Label Encoding
D. Dummy Coding

Answer: B


18. Which of the following is an unsupervised learning technique?

A. Logistic Regression
B. Support Vector Machine
C. Hierarchical Clustering
D. Random Forest

Answer: C


19. Which statement about Naïve Bayes is TRUE?

A. It assumes attributes are mutually exclusive.
B. It assumes conditional independence among predictors.
C. It requires numerical attributes only.
D. It cannot classify multiclass problems.

Answer: B


20. Which kernel allows Support Vector Machines to classify non-linearly separable data?

A. Identity Kernel
B. Polynomial or RBF Kernel
C. Mean Kernel
D. Binary Kernel

Answer: B


21. Which ensemble learning technique primarily aims to reduce variance?

A. Boosting
B. Bagging
C. Gradient Descent
D. PCA

Answer: B


22. In Random Forest, randomness is introduced by:

A. Using different distance metrics
B. Selecting random subsets of data and features
C. Changing class labels
D. Randomly removing decision trees

Answer: B


23. Which data quality issue occurs when the same real-world entity appears multiple times with different identifiers?

A. Missing Values
B. Duplicate Records
C. Inconsistent Encoding
D. Data Drift

Answer: B


24. Which OLAP operation transforms detailed data into summarized data?

A. Drill-down
B. Slice
C. Roll-up
D. Pivot

Answer: C


25. Which statement about Hadoop MapReduce is correct?

A. Mapping aggregates final results.
B. Reduce executes before Map.
C. Mapping processes input into intermediate key-value pairs before reduction.
D. MapReduce requires a relational database.

Answer: C


26. Which of the following is not one of the "3 Vs" originally associated with Big Data?

A. Volume
B. Velocity
C. Variety
D. Validity

Answer: D


27. In Apache Spark, which component is primarily responsible for in-memory distributed computation?

A. HDFS
B. RDD (Resilient Distributed Dataset)
C. Hive
D. Sqoop

Answer: B


28. Which ETL operation is responsible for resolving inconsistent date formats before loading data into a warehouse?

A. Extraction
B. Transformation
C. Loading
D. Replication

Answer: B


29. Which normalization method scales data into the interval [0,1]?

A. Z-score Normalization
B. Decimal Scaling
C. Min-Max Normalization
D. Mean Centering

Answer: C


30. A model achieves 99% accuracy on a dataset where 99% of observations belong to one class. This is an example of:

A. High bias
B. Data leakage
C. Misleading accuracy due to class imbalance
D. Underfitting

Answer: C


31. Which metric combines Precision and Recall into a single harmonic mean?

A. ROC Score
B. Specificity
C. F1 Score
D. Lift

Answer: C


32. ROC curves plot:

A. Precision against Recall
B. Recall against Accuracy
C. True Positive Rate against False Positive Rate
D. Sensitivity against Specificity

Answer: C


33. An AUC (Area Under Curve) value of 0.5 indicates:

A. Perfect classifier
B. Random guessing
C. Excellent classifier
D. Severe overfitting

Answer: B


34. Which statistical measure is least affected by extreme outliers?

A. Mean
B. Variance
C. Median
D. Standard Deviation

Answer: C


35. Which technique is commonly used to detect multicollinearity among predictor variables?

A. Silhouette Score
B. Variance Inflation Factor (VIF)
C. Lift Ratio
D. Entropy

Answer: B


36. Which clustering evaluation metric does not require ground-truth labels?

A. Accuracy
B. Silhouette Coefficient
C. Precision
D. Recall

Answer: B


37. In DBSCAN, points that are neither core points nor border points are called:

A. Seed Points
B. Noise Points
C. Pivot Points
D. Anchor Points

Answer: B


38. Which of the following algorithms is deterministic (produces the same result every run with identical data)?

A. Standard K-Means
B. Random Forest
C. Hierarchical Agglomerative Clustering
D. Bagging

Answer: C


39. Which distance measure is most suitable when variables have different units without prior normalization?

A. Euclidean Distance
B. Manhattan Distance
C. Mahalanobis Distance
D. Hamming Distance

Answer: C


40. Which technique reduces overfitting by penalizing large regression coefficients?

A. Bootstrap Sampling
B. Regularization
C. One-Hot Encoding
D. Cross Join

Answer: B


41. LASSO regression differs from Ridge regression because it:

A. Uses no penalty term
B. Can shrink some coefficients exactly to zero
C. Always produces larger coefficients
D. Works only for classification

Answer: B


42. Which assumption is essential for Linear Regression but not for Decision Trees?

A. Presence of categorical variables
B. Linear relationship between predictors and response
C. Missing values
D. Binary output

Answer: B


43. In time series analysis, autocorrelation refers to:

A. Correlation between independent variables
B. Correlation of a variable with its own past values
C. Correlation between training and testing data
D. Correlation between clusters

Answer: B


44. Which forecasting model explicitly includes trend, seasonality, and error components?

A. K-Means
B. Holt-Winters Exponential Smoothing
C. Naïve Bayes
D. Apriori

Answer: B


45. Which SQL operation is most commonly used during ETL to combine records from multiple tables?

A. DROP
B. ALTER
C. JOIN
D. TRUNCATE

Answer: C


46. Which NoSQL database type is best suited for storing graph relationships such as social networks?

A. Key-Value Store
B. Column Family Database
C. Graph Database
D. Document Store

Answer: C


47. Which Apache Hadoop ecosystem component provides SQL-like querying over distributed data?

A. Pig
B. Hive
C. Flume
D. Oozie

Answer: B


48. Which of the following is an example of descriptive analytics?

A. Predicting customer churn
B. Estimating next year's revenue
C. Summarizing last quarter's sales by region
D. Optimizing delivery routes

Answer: C


49. Which type of analytics answers the question "What should we do?"

A. Descriptive Analytics
B. Diagnostic Analytics
C. Predictive Analytics
D. Prescriptive Analytics

Answer: D


50. Which statement about data governance is most accurate?

A. It deals only with database security.
B. It focuses only on backup and recovery.
C. It establishes policies, standards, ownership, and accountability for data assets.
D. It is synonymous with data mining.

Answer: C

51. In association rule mining, a Lift value greater than 1 indicates:

A. The antecedent and consequent are negatively correlated.
B. The antecedent and consequent occur independently.
C. The occurrence of the antecedent increases the likelihood of the consequent.
D. The rule has zero confidence.

Answer: C


52. Which association rule measure is symmetric, unlike confidence?

A. Support
B. Lift
C. Confidence
D. Conviction

Answer: B


53. Which algorithm is specifically designed for anomaly detection by randomly partitioning data?

A. K-Means
B. Isolation Forest
C. Apriori
D. FP-Growth

Answer: B


54. Local Outlier Factor (LOF) identifies anomalies based on:

A. Global mean deviation
B. Density relative to neighboring observations
C. Euclidean distance from origin
D. Cluster centroids only

Answer: B


55. Concept drift refers to:

A. Missing values in a dataset.
B. Gradual or sudden changes in the statistical properties of data over time.
C. Loss of database indexes.
D. Duplicate records in a warehouse.

Answer: B


56. Which technique is commonly used to handle highly imbalanced classification datasets?

A. PCA
B. SMOTE
C. K-Means
D. Min-Max Scaling

Answer: B


57. SMOTE generates:

A. Random noise variables.
B. Synthetic minority class samples.
C. Additional majority class records.
D. Duplicate observations only.

Answer: B


58. Which ensemble algorithm builds trees sequentially, where each new tree attempts to correct previous errors?

A. Random Forest
B. Bagging
C. Gradient Boosting
D. Extra Trees

Answer: C


59. XGBoost improves Gradient Boosting mainly through:

A. Removing regularization.
B. Parallelization and regularization techniques.
C. Eliminating decision trees.
D. Using only categorical variables.

Answer: B


60. Which impurity measure can never be negative?

A. Entropy
B. Gini Index
C. Variance
D. All of the above

Answer: D


61. In information theory, entropy becomes zero when:

A. All classes occur with equal probability.
B. Every class has exactly two observations.
C. The dataset contains only one class.
D. The dataset has missing values.

Answer: C


62. Which theorem forms the mathematical foundation of the Naïve Bayes classifier?

A. Central Limit Theorem
B. Bayes' Theorem
C. Markov Theorem
D. Chebyshev's Theorem

Answer: B


63. Which NLP technique converts text into numerical vectors based on word frequencies?

A. PCA
B. TF-IDF
C. One-Hot Encoding of Classes
D. Z-score Scaling

Answer: B


64. In TF-IDF, the IDF component gives higher weight to words that:

A. Appear in every document.
B. Are rare across documents.
C. Are longest in length.
D. Contain numerical values.

Answer: B


65. Which web mining category focuses on analyzing hyperlink structures?

A. Web Usage Mining
B. Web Structure Mining
C. Web Content Mining
D. Stream Mining

Answer: B


66. Clickstream analysis is primarily associated with:

A. Web Usage Mining
B. Data Warehousing
C. OLTP Systems
D. Image Mining

Answer: A


67. Stream mining algorithms must primarily address:

A. Unlimited storage availability.
B. Infinite processing time.
C. Continuous arrival of data with limited memory.
D. Static datasets only.

Answer: C


68. Which property best distinguishes a Data Lake from a traditional Data Warehouse?

A. Data Lake stores only structured data.
B. Data Warehouse stores only unstructured data.
C. Data Lake can store structured, semi-structured, and unstructured data in raw form.
D. Data Warehouse requires Hadoop.

Answer: C


69. Which privacy model extends k-anonymity by ensuring diversity in sensitive attributes?

A. PCA
B. l-diversity
C. FP-Growth
D. ROC

Answer: B


70. Differential Privacy primarily aims to:

A. Compress datasets.
B. Protect individual information by adding controlled statistical noise.
C. Increase classification accuracy.
D. Improve clustering quality.

Answer: B


71. Which sampling method gives every population member an equal probability of selection?

A. Stratified Sampling
B. Cluster Sampling
C. Simple Random Sampling
D. Systematic Sampling

Answer: C


72. Which statistical test is generally used to compare the means of two independent samples?

A. Chi-Square Test
B. t-Test
C. ANOVA
D. Mann-Whitney Test

Answer: B


73. Which hypothesis error occurs when a true null hypothesis is rejected?

A. Type II Error
B. Sampling Error
C. Type I Error
D. Standard Error

Answer: C


74. A p-value less than the chosen significance level (α) indicates:

A. Accept the null hypothesis.
B. Reject the null hypothesis.
C. Increase sample size immediately.
D. The experiment is invalid.

Answer: B


75. Which statement about Explainable AI (XAI) is most accurate?

A. It aims to replace machine learning models with SQL queries.
B. It focuses on making AI model decisions understandable to humans.
C. It eliminates the need for model validation.
D. It guarantees 100% prediction accuracy.

Answer: B

76. Which OLAP operation creates a smaller cube by selecting a single value for one of the dimensions?

A. Roll-up
B. Drill-down
C. Slice
D. Pivot

Answer: C


77. Which OLAP operation changes the dimensional orientation of a data cube to provide an alternative presentation of data?

A. Drill-down
B. Pivot (Rotate)
C. Roll-up
D. Slice

Answer: B


78. Which data cube computation strategy computes only the cuboids that are actually required, thereby reducing storage and computation?

A. Full Materialization
B. No Materialization
C. Partial Materialization
D. Complete Aggregation

Answer: C


79. Which optimization technique significantly improves the efficiency of the Apriori algorithm?

A. Candidate pruning using the downward closure property
B. Increasing the minimum support threshold after every iteration
C. Sorting transactions alphabetically
D. Using Euclidean distance

Answer: A


80. Sequential Pattern Mining differs from Association Rule Mining because it considers:

A. Only numerical attributes
B. Temporal or ordered relationships among events
C. Only binary data
D. Data compression techniques

Answer: B


81. Which feature selection method evaluates variables independently of any machine learning algorithm?

A. Wrapper Method
B. Embedded Method
C. Filter Method
D. Ensemble Method

Answer: C


82. In Principal Component Analysis (PCA), the first principal component is the one that:

A. Has the smallest variance
B. Is most correlated with the target variable
C. Explains the maximum variance in the data
D. Contains only categorical variables

Answer: C


83. Which matrix is decomposed into eigenvalues and eigenvectors during Principal Component Analysis?

A. Identity Matrix
B. Covariance Matrix (or Correlation Matrix after standardization)
C. Confusion Matrix
D. Transition Matrix

Answer: B


84. Covariance differs from correlation because covariance:

A. Is always between –1 and +1
B. Is independent of the units of measurement
C. Depends on the scale (units) of the variables
D. Can never be negative

Answer: C


85. Which Apache Spark component is responsible for coordinating the execution of an application?

A. Worker Node
B. Driver Program
C. DataNode
D. NameNode

Answer: B


86. In the Hadoop ecosystem, YARN is primarily responsible for:

A. Data compression
B. Resource management and job scheduling
C. SQL querying
D. Data visualization

Answer: B


87. According to the CAP Theorem, a distributed database can guarantee at most which two of the following three properties simultaneously?

A. Capacity, Availability, Performance
B. Consistency, Availability, Partition Tolerance
C. Consistency, Accuracy, Partition Tolerance
D. Availability, Performance, Security

Answer: B


88. Which property is associated with NoSQL databases under the BASE model rather than the ACID model?

A. Atomicity
B. Immediate Consistency
C. Eventual Consistency
D. Isolation

Answer: C


89. Which SQL window function assigns the same rank to tied rows while leaving gaps in subsequent ranks?

A. ROW_NUMBER()
B. DENSE_RANK()
C. RANK()
D. NTILE()

Answer: C


90. Data lineage refers to:

A. The genealogy of employees in an organization
B. The lifecycle of data, including its origin, transformations, and movement
C. The structure of a database schema only
D. Backup history of a database

Answer: B


91. Metadata is best described as:

A. Duplicate data
B. Data about data
C. Temporary data
D. Backup data

Answer: B


92. Which phase of MLOps focuses on continuously tracking model performance after deployment?

A. Feature Engineering
B. Model Monitoring
C. Data Cleaning
D. Data Integration

Answer: B


93. Model drift occurs when:

A. The programming language changes
B. The statistical relationship between input data and target variable changes over time
C. Storage capacity decreases
D. The model file becomes corrupted

Answer: B


94. Which statement best describes the Bias–Variance Tradeoff?

A. Reducing bias always reduces variance.
B. Increasing model complexity generally reduces bias but may increase variance.
C. Bias and variance are unrelated.
D. High variance always improves generalization.

Answer: B


95. The primary objective of an A/B test is to:

A. Compress large datasets
B. Compare two alternatives using statistical evidence
C. Remove missing values
D. Build classification trees

Answer: B


96. A Bayesian Network is best described as:

A. A relational database model
B. A probabilistic graphical model representing conditional dependencies among variables
C. A clustering algorithm
D. A distributed file system

Answer: B


97. Which probability theorem is commonly used to compute the probability of an event by partitioning the sample space into mutually exclusive cases?

A. Bayes' Theorem
B. Law of Total Probability
C. Chebyshev's Inequality
D. Markov Property

Answer: B


98. Which measure is most appropriate for evaluating regression models?

A. Accuracy
B. F1-Score
C. Mean Squared Error (MSE)
D. Precision

Answer: C


99. Which statement about feature engineering is correct?

A. It is useful only for deep learning models.
B. Creating meaningful features from raw data can significantly improve model performance.
C. It always increases the number of features.
D. It is performed only after model deployment.

Answer: B


100. Which of the following best summarizes the goal of data mining?

A. To store data efficiently in databases.
B. To retrieve predefined reports from databases.
C. To discover valid, novel, useful, and understandable patterns from large datasets for decision-making.
D. To replace statistical analysis entirely.

Answer: C


1 $type={blogger}:

akgvg accounting service said...


AKGVG is the leading audit firm in India. We are a team of proficient and dedicated chartered accountants based in New Delhi as well as other major cities in India. The Audit Services are tailored to your specific needs, and combined as relevant.


audit firms in Delhi