tutorbin

data mining homework help

Boost your journey with 24/7 access to skilled experts, offering unmatched data mining homework help

tutorbin

Trusted by 1.1 M+ Happy Students

WhatsApp Support

Get Instant
Online Homework Help
via WhatsApp

Get instant homework help from top tutors—just a WhatsApp message away. 24/7 hw help support for all your academic needs!

A
S
M
R
★★★★★
2M+ students trust TutorBin
Your WhatsApp Number
phone
or
⚡ Instant reply
🔒 100% private
👨‍🏫 Top tutors
🌍 All subjects
*Get instant homework help from top tutors—just a WhatsApp message away. 24/7 support for all your academic needs!
2M+ Students Helped24/7 Live SupportExpert TutorsAll Subjects CoveredInstant Response100% ConfidentialTop Rated ServiceMoney-back Guarantee2M+ Students Helped24/7 Live SupportExpert TutorsAll Subjects CoveredInstant Response100% ConfidentialTop Rated ServiceMoney-back Guarantee

Recently Asked data mining Questions

Expert help when you need it
  • Q1: Instructions for submitting the solution 1. Submit the source code file for every programming assignment like .c file, .java file. 2. If any dataset is used then you need to share all the files with solution in a single zip file. 3. Share the readme file in which you have to include the system specification, required software and the execution instructions. 4. Share the output of the programs. If it is a single program then you can submit the screenshot, if multiple files are included then please share the screen recording. Below are some steps that need to be covered in screen recording. ❖ Show the complete solution according to the instructions. ❖ Need to cover all the compilation steps. Show the complete output of program/project. Show all pass test cases if given in assignment. 5. Always add proper comments in code. If any specific package is used then please mention it in comment./n UCF-CECS Using Python & a screen scraper to extract data from Wikipedia Homework Assignment 3 (hw03) April 1, 2024 1 Objectives The goal of this homework assignment is to develop a solution to extract (screen scrape) the SpaceX Falcon 9 Block 5 launch records from Wikipedia. It is important to note that this website is being actively & substantially modified. So a stable version of the Wikipedia page for this assignment has been supplied. Use the supplied file is to develop a solution to extract (screen scrape) data for the SpaceX Falcon 9/Heavy Launches using the Block 5 engines. This is described in further detail below. 1.1 Specific data There are several data in the webpage elements to be extracted in this assignment. They are as follows: • All Block 5 engines' launch history ((may vary from one to many launches). - Launch identified by launch number (Fx-DDD where x is either a 9 for Falcon 9, or H for a Falcon Heavy and DDD is a 3 decimal digit) - - Launch date in DD-Month-Year format Turnaround time in days CIS4340-McAlpin HW 03 1 • In this assignment it is important to note the Falcon 9 launches are a single launch booster using a Block 5 engine, and Falcon Heavy launches have three Block 5 engines. These objectives will be met and demonstrated in the exercises specified later in the assignment. 1.2 Collected data As discussed earlier, the Falcon 9 wikipedia page is currently being substantially revised. This page is currently being split. After a discussion, consensus to split this page into List of Falcon 9 and Falcon Heavy launches (2020-2021) was found. You can help implement the split by following the instructions at Help:Splitting and the resolution on the discussion. Process started in March 2024. For this reason, the file (Falcon9first-stageBoosters.html) has been supplied for this assignment. It is in the Webcourses assignment page. 2 CIS4340-McAlpin HW 03 S/Nial Type Launches Launch date (UTC)[5] Falcon 9 block 5 first-stage boosters Expended, Destroyed, or Officially Retired) Flight No. Turnaround Payload cl 11 May 2018 [b] F9-054 time Bangabandhu-188] 7 August 2018 E9-06088 days Telkom-4 Merah Putih 891 B1046 E9 3 December 2018 P9-064 118 days 19 January 2020911 F9-079412 days SHERPA (SSO-A)[88][90] Dragon C205 (In-Flight Abort Testy 921 22 July 2018 F9-058 - Telstar 19y1931 B1047 9 15 November 2018 F9-063116 days Es'hail 2241 Launch Landing (pad) (location) Success Success (39A) (OCISLY) Success Success (40) (OCISLY) Success (4E) Success (39A) Success (40) Success (39A) Success Status Expended Success (JRTI) No attempt (OCISLY) Success Expended (OCISLY) 6 August 20191951 E9-074263 days AMOS-17 No attempt 271 (40) B1024's history in html: <tr id="B1024"> <td>B1024 </td> Figure 1.1: The first 7 rows of Block 5 engines' data <td><a href="/wiki/Falcon_9_Full_Thrust" title="Falcon 9 Full Thrust">FT</a> </td> <td><span data-sort-value="000000002016-06-15-0000" style="white-space: nowrap">15 June 2016</span> </td> <td><a href="/wiki/Falcon_9_flight_26" title="Falcon 9 flight 26">F9-026</a> </td> <td data-sort-value="" style="background: #ececec%; color: #2C2C2C; vertical-align: middle; text-align: center;" class="table-na">âĂŤ </td> <td><a href="/wiki/ABS-2A" title="ABS-2A">ABS-2A</a> / <a href="/wiki/Eutelsat_117_West_B" class="mw-redirect" title="Eutelsat 117 West B">Eutelsat 117 West B</a> </td> <td style="background: #9EFF9E; vertical-align: middle; text-align: center;" class="table-success">Success<br />(40) </td> <td style="background: #FFC7C7%3B vertical-align: middle; text-align: center;" class="table-failure">Failure </td> <td>Destroyed<sup id="cite_ref-SFN-2017-06-15_44-0" class="reference"><a href="#cite_note-SFN-2017-06-15-44">&#91;40&#93;</a></sup> </td></tr> The rows, <tr>, and cells by column, <td> tags support configuration for the number of rows in each column. Those tags are navigable in the scraping code, often indexable too. CIS4340-McAlpin HW 03 3 S/Nlal Type Launches uch date TC151 Flight No. Turnaround F9-159 daigne [156] Launch Landing Starlinkpl L19) Success (pad) (location) (40) (OCISLY)[157] Status B1066 EH core 1 1 November 2022 FH-004 USSF-44 B1068 FH core 1331 1 1 May 2023[131] FH-006 - ViaSat-3 Americas[131] B1070 FH core 1 15 January 2023[158] FH-005 - USSF-67 Success (39A) Success (39A) Success (39A) No attempt Expended No attempt[132] Expended No attempt Expended B1074 FH core 1 29 July 2023 FH-007 - Jupiter-3 (EchoStar-24) Success (39A) Success No attempt Expended B1079 FH core 1 13 October 2023 FH-008 Psyche 1591 No attempt Expended (39A) B1084 FH core 1 29 December 2023 FH-009 USSF-52 (Boeing X-37B OTV-7) Success (39A) No attempt Expended 1.3 Programs 1.3.1 Extraction Figure 1.2: The last 6 rows of Block 5 engines' data The following data needs to be extracted using a screen scraper applied to the HTML file, Falcon9first-stageBoosters.html, which is supplied via Webcourses. 1. The Block 5 engine number. 2. The Flight number. 3. The Flight type a) F9 for Falcon 9 b) FH for Falcon Heavy 4. The launch date, in the YYYY-MM-DD format. 5. The launch pad. 6. The landing location, typically an acronym, sometimes also identified as No attempt. 7. The Turnaround time, in days. 8. The engine's status: a) Expended b) Destroyed c) Lost at sea d) Returned to service 9. The total number of launches for this engine. These data elements should be output in the order shown above, to STDOUT. Each element should be separated by a comma, thereby building a CSV. This Python program should be named, Block5Extract.py. CIS4340-McAlpin HW 03 4 Wondering how to run the program and capture the output? - - xyz$python3 Block5Extract.py > Block5.csv 1.3.2 Reports * This command prompt executes the Python program using the Falcon9first-stageBoosters.html as input and redirects the output from STDOUT to the file named Block5.csv. * Make sure Falcon9first-stageBoosters.html is in the same directory as the code. # Title 1 f9only 2 fHonly 3 fHpairs 4 5 6 Table 1.1: Report names Description Only Falcon 9 launches Only Falcon Heavy launches The three engines used for each Falcon Heavy launch longestTurnaround | The longest turnaround for a Block 5 engine fastestTurnaround mostLaunches Notes: The fastest turnaround for a Block 5 engine The most number of launches for a Block 5 engine Use the Title as shown above for both the program name, i.e. #1 would be f9only .py and the output to be redirected to the filename f9only.txt. Both the program and the output file for each of the 6 programs/reports will be submitted to Webcourses. 1.3.3 Submission instructions You must submit this assignment in Webcourses as file uploads. It is preferred to ZIP your submissions. The submitted programs are as follows: 1. f9only 2. fHonly 3. fHpairs 4. longestTurnaround 5. fastestTurnaround 6. mostLaunches 7. Block5Extract CIS4340-McAlpin HW 03 5/nSee Answer
  • Q2: 1 Complete the following steps and then answer the exercise questions below. Step 1. Import the training and scoring data sets for this exercise into data frames in RStudio. 2 3 Step 2. Load the R library required to create a logistic regression model. Step 3. Create a logistic regression model to predict RenewedSubscription. Do not include PatronID as an independent variable. Coerce the dependent variable to be treated as a factor. Use the summary() function to inspect your independent variables' p-values. Do not remove any independent variables from the model. Step 4. Using a subset() command, remove observations, if any, from the scoring data set where one or more attributes exceed the range established in the training data set. For example, the range for DifferentUsers in the training data set is 2 to 12. If any observations in the scoring data set have DifferentUsers values below 2 or above 12, remove them. Check all attributes to ensure all scoring observations are within ranges established by the training data. Step 5. Using the predict() function, apply your logistic regression model to the scoring data. Make sure the type of prediction you generate is the model's "response." Step 6. Combine the predictions with the scoring data into a new data frame. View the data frame and answer the following questions. Which attribute is the single poorest predictor of season ticket renewal? ☐ PricePerTicket AvgMinutes BeforeCurtain ConcessionVouchers PerformancesAttended Of the first-year season ticket patrons in the scoring data set, how many are predicted by the logistic regression model to renew their subscriptions? 68 80 65 148 Considering all properties of the logistic regression model, which attribute is a better predictor of season ticket renewal? PricePerTicket ☐ NumberOfTickets PricePerTicket and NumberOfTickets have the exact same predictive strength in this model. Neither PricePerTicket nor NumberOfTickets have predictive strength in this model. If you wished to test the accuracy of your logistic regression model in R, to which data set would you apply the predict() function? The test data The training data The scoring data The validation data 5 How many "No" predictions have a post-probability confidence percent higher than 95%? ☐ 9 80 ☐ 68 71 6 Complete the following steps and then answer the exercise questions below. Step 1. Import the training and scoring data sets for this exercise into data frames in RStudio. Step 2. Load the R libraries required to create a decision tree model and to visualize the model graphically. Step 3. Create a decision tree model to predict InsuranceCategory. Do not include CustomerID as an independent variable. Step 4. Use the summary() function to inspect your decision tree's properties. Do not remove any independent variables from the model. Step 5. Use the predict() function to apply your decision tree model to the scoring data frame to predict InsuranceCategory classes for each observation. Store these predictions in a data frame with a relevant name. Step 6. Use the predict() function to apply your decision tree model to the scoring data frame to predict Insurance Category confidence percentages (post-probabilities) for each observation. Store these predictions in a data frame with a relevant name. Step 7. Create a data frame with a relevant name that contains the Step 5 class predictions, the Step 6 confidence percentages, and the scoring data. Step 8. Enlarge the Plots pane of RStudio and then create a visual depiction of your decision tree. Using the visual depiction (plot) of the decision tree model, which attribute is the first, best predictor of insurance category? NumberOfClaims Age LatePayments ☐ AtFaultAccidents 7 Using the visual depiction (plot) of the decision tree model, what percent of the training observations have one or more at-fault accidents but no comprehensive claims on their insurance policies? 55% 19% 38% 7% 8 The summary() description of the decision tree model shows seven of the eight independent variables under Variable Importance. Which of the seven listed is the least important? MovingViolations ☐ Gender Age CompClaims 9 The summary() description of the decision tree model shows seven of the eight independent variables under Variable Importance. Which of the independent variables in the R model is deemed unimportant? Age CompClaims MaritalStatus U customerID K K K K 10 In the visual depiction (plot) of the decision tree model, only one leaf of the tree leads to a prediction of High Risk-Do Not Insure. Not all of the training observations that follow that branch of the tree are classified in that category, however. What percent of those observations are actually classified into the Potentially High Risk category? 36% 0% 62% 2% 11 The data sets used for this end-of-chapter exercise will need to be normalized in order to create a neural network that can produce reliable predictions. This will be accomplished using the scale() function in R, which has not been covered in the text. The modifications to ensure that nnet() produces usable results will therefore be prescribed in the following steps. Complete each step and then answer the exercise questions below. Step 1. Import the training and scoring data sets for this exercise into data frames in RStudio. Ensure that the training data frame's name is ch11Train and the scoring data frame's name is ch11Score. Step 2. Load the R library required to create a neural network model. Step 3. Use the following two commands to create data frames containing the normalized independent variable values for the training and scoring data sets. ch11TrainNorm <- data.frame(scale(ch11Train[2:8])) ch11ScoreNorm <- data.frame(scale(ch11Score [2:8])) Step 4. The independent variable values that were normalized using the scale() function in step 3 will be used to train the neural network model. Issue the command: attach(ch11TrainNorm) Step 5. Set the seed value to 43. Step 6. You will train a neural network model to predict the Credit Risk dependent variable in the ch11Train data frame, using the independent variables in the normalized ch11TrainNorm data frame. The hidden layer's size attribute is set to 7, which uses the generally accepted formula: ((7 independent variables + 5 dependent variable levels)/2) + 1; which is (12/2) + 1 = 7 Issue the command: ch11NNModel <- nnet(as.factor(ch11 Train$Credit Risk) - Credit Score+Late_Payments+Months In Job+Debt Income Ratio+Loan_Amt+Liquid Assets+Num_Credit Lines, data=ch11TrainNorm, size=7, maxit=1000) Step 7. Use the predict() function to apply your neural network model (ch11NNModel) to the normalized scoring data frame (ch11ScoreNorm) to predict Credit Risk classes for each observation. Store these predictions in a data frame with a relevant name. Step 8. Use the predict() function to apply your neural network model (ch11NNModel) to the normalized scoring data frame (ch11ScoreNorm) to predict Credit Risk confidence percentages (post-probabilities) for each observation. Store these predictions in a data frame with a relevant name. Step 9. Create a data frame with a relevant name that contains the step 7 class predictions, the step 8 confidence percentages, and the unnormalized scoring data (ch11Score). View this data frame in RStudio, then answer the following questions. Using the neural network's predictions as generated using the steps outlined in this exercise, how many loan applicants will have their loans denied? 79 11 23 12 Normalizing the data has resulted in most confidence percentages being 100% in this exercise's predictions. However, not all predictions have 100% confidence. How many predictions of high credit risk in this neural network's results are not 100%? 162 2 12 79 77 K 13 Using the neural network's predictions as generated using the steps outlined in this exercise, and assuming that loan officers can automatically approve all loans predicted to be low or very low risk, how many loans will be automatically approved by loan officers? 126 114 142 16 14 Assume the bank has agreed to approve the loans for the following applicants: 931184, 937005, 451482, and 597325. Based on the model's predictions and confidence percentages, which applicant do you expect will be offered the least favorable terms and the highest interest rate? 597325 931184 937005 15 451482 Based on this model's predictions, the bank is 71% confident that Applicant ID 311882 is high risk. Using the predictions, how confident is the bank that this applicant may actually pose only a moderate risk in lending? 28.9% 15.6% 100% 71.1%See Answer
  • Q3:Part 2: You are a data analyst and have been tasked with determining whether residents of the state of Montana should expect the economic outlook for their state to be better than it was a year ago. You will use the "Poll" dataset to help you make this determination. The dataset contains demographic information from respondents along with information about their financial status. Additionally, the website includes the coding for dataset variables. For the STAT field, "0" indicates the state economic outlook will be better than a year ago, while a "1" indicates that the state economic outlook will not be better than it was a year ago. Provide a histogram of the STAT field. Then, perform an association analysis using the Apriori algorithm; use the KNIME Association Rule Learner node to perform the analysis. Experiment with the minimum confidence and minimum support settings in this node in order derive meaningful association rules. Note that we are only interested in rules in which STAT is the consequent. Note that a "Create Collection Column" node needs be placed between the Excel Reader node and the Association Rule Learner node. In the Create Collection Column node settings, ensure that the following boxes are checked: Create a collection of type 'set' (doesn't store duplicate values) ignore missing values Remove aggregated columns from table Also, ensure that all fields in the dataset are shown in the "Include" portion of the dialog box./nUse a Table node to see the derived association rules. In a 250-500-word summary, address the following: 1. Explain your approach to the problem. 2. Include screenshots of the histogram results of the STAT field and the Table node screen results of the derived association rules. Ensure that the rules are sorted by Lift, in descending order. 3. Interpret the relevant derived association rules by describing how you would explain the rules to someone who is unfamiliar with analytics techniques. 4. Summarize your overall conclusion about the data based upon the results of the analysis. Note that you are required to submit the completed KNIME *.knwf file to your instructor. Specifically, export your KNIME model to a KNIME workflow file. To perform this task in KNIME, ensure that your KNIME model is active (i.e., displayed). Then, go to File -> Export KNIME Workflow. In the "Destination workflow file name (.knwf)" area, browse to a specific location on your computer. Click "Save" and then click "Finish." Submit the Excel file, Word document, and the KNIME *.knwf workflow file.See Answer
  • Q4:Part 1: Use the "Weather" dataset to address the questions below. 1. Using +-4, generate the frequent 2-itemsets in the "Weather" dataset. 2. Using 75% minimum confidence and 20% minimum support, generate one-antecedent association rules for predicting "Play" (i.e., Play= "Yes") using the "Weather" dataset. List the support and confidence along with each rule. Note that Play = "Yes" means the game can be played indoors and Play = "No" means the game cannot be played indoors. Note that Rule Support and Confidence calculations are needed. In addition, be sure to specifically explain (in understandable terms) the specific association rules generated.See Answer
  • Q5: Instructions Data Mining The purpose of this assignment is to demonstrate understanding and application of the k-Nearest Neighbor algorithm associated with classification. You have been asked to build a classification model to help predict machine failure. Using the "Failure Rate" dataset, build a classification model using the k-Nearest Neighbor technique in KNIME. Then, using the model, predict whether machines will failure for 50 records of input data. Training/Test Model Use "hours_run" and "avg_hours_between_maint" as the input variables. Use "failure" as the target variable (note that 0: = no failure, and 1 = failure). For the Excel Reader node, exclude the "model_version" field from the data import operation. Use a Normalizer node to normalize the "hours_run" and "avg_hours_between_maint" fields to be between 0 and 1. Use a Partition node with an 70/20 partition (i.e., 70% Training; 20% Test) for records 1-300. Then, create the model using the Training data from the Partition Node. Use the k-Nearest Neighbor node. Attach a Scorer node to the k-Nearest Neighbor node in order to evaluate the model's accuracy (note that this node only evaluates the results of the n=90 Test data). Attach a Table node to the second output port of the Scorer node. Take note of the overall accuracy value. Run the model for k values from 3 to 6. Select the base k value based on the highest accuracy value. Predictions on the n=50 Data After running the Training/Test model and ascertaining the optimal k value, create another workflow (in the same KNIME file) to predict machine failure for the rest of the records on the dataset, i.e., for records 301-350. Setup the workflow as shown in the attached "K-Nearest Neighbor Algorithm, Prediction Workflow" document. Attach a Table node to the k-Nearest Neighbor node to view the predicted results. Sort by the appropriate column to show those records whose machines are predicted to fail at the top of the list. Also, attach an Excel Writer node to the k-Nearest Neighbor node to export the prediction results in an Excel file. In a 250-word document, provide the following information. Assume you are providing this information to an audience that has limited knowledge of data mining concepts. 1. Summarize your approach to the problem. 2. Clearly state the optimal k value for the model. 3. Screenshot the results from the second output port of the Scorer node. 4. Screenshot the predicted results, sorted by appropriate column to show at the top of the list those records whose machines are predicted to fail. 5. Include a conclusion based on the results of the analysis. Specify which machines in records 301-350 are predicted to fail. Speculate on why these particular records are predicted to fail. Note that you are required to submit the completed KNIME *.knwf file to your instructor. Specifically, export your KNIME model to a KNIME workflow file. To perform this task in KNIME, ensure that your KNIME model is active (i.e., displayed). Then, go to File -> Export KNIME Workflow. In the "Destination workflow file name (.knwf)" area, browse to a specific location on your computer. Click "Save" and then click "Finish."/n Instructions Using specified data files, chapter example files, and templates from the "Topic 4 Student Data, Template, and Example Files" resource, complete Chapter 13 Problems 20, 26, 28, 50, and 52 in the textbook. Use MAPE (mean absolute percentage error) to evaluate the forecasting performance for each problem. Use the Palisade Decision Tools Excel software to complete these problems where requested and applicable. To receive full credit on the assignment, ensure that the Excel files include the associated cell formulas if formulas are used or Excel-generated output based on the nature of the analyses. Place each problem in its own Excel file. Ensure that your first and last name are in your Excel filenames. THEN The purpose of this assignment is to conduct analyses and present your findings and supporting documentation in a professional PowerPoint presentation designed to summarize the information for senior leadership within the organization. Assume that you are delivering this presentation to the senior leadership in an organization. Therefore, please be sure to create a professional presentation. Begin by reading the "13.2 Forecasting Overhead at Wagner Printers" case, found at the end of Chapter 13 in the textbook. For the case, you will perform a multiple regression analysis. You can perform additional analyses on each data set to gain greater insight into the data set. You must be able to justify each of the approaches and methods you selected for analyzing the data sets. Use the Palisade DecisionTools Excel software to perform the regression analysis. Evaluate the regression model by performing and responding to all parts of the "Multiple Regression Analysis Checklist." Use the "BIT-435-RS-Predictive Case Template and Support Files" to complete the assignment and submit answers. Prior to submission, rename the file to include your first name and last name in the filename. You will submit the completed template file along with the PowerPoint presentation. Results of each analysis must be included in your presentation. The use of graphs, charts and supporting data, and spreadsheets is encouraged. Interpret the results of each analysis and draw general conclusions from the results. Make recommendations for the organization and address the organizational challenges that may be encountered based upon your recommendations. The PowerPoint presentation should include the following information: 1. Introduction and case background. 2. Objectives for each analysis. 3. Approach or method of analysis for each data set and justification for selecting the approach or method. Results of each analysis. 4. 5. Supporting graphs, charts, data, and spreadsheets for each analysis. 6. Interpretation of the results for each analysis. 7. General conclusion of each analysis and recommendation to the organization, including addressing organizational challenges that may be encountered based upon the recommendation. 8. In the Speaker Notes section of each slide, include your talking points. This information should align to the results of your analyses and be supported in the accompanying Excel files. In addition to your PowerPoint file, submit the completed template file that contains the supporting Excel files showing all data analyses performed. Submission of your Excel files is required to obtain full credit for this assignment./n 50. The file P13_50.xlsx contains five years of monthly data for a company. The first variable is Time (1-60). The second variable, Sales1, has data on sales of a product. Note that Sales1 increases linearly through- out the period, with only a minor amount of noise. (The third variable, Sales2, will be used in the next problem.) For this problem, use the Sales1 variable to see how the following forecasting methods are able to track a linear trend. a. Forecast this series with the moving averages method with various spans such as 3, 6, and 12. What can you conclude? b. Forecast this series with simple exponential smoothing with various smoothing constants such as 0.1, 0.3, 0.5, and 0.7. What can you conclude? c. Repeat part b with Holt's method, again for vari- ous smoothing constants. Can you do much better than in parts a and b? 52. The file P13_52.xlsx contains data on a motel chain's revenue and advertising. a. Use these data and multiple regression to make pre- dictions of the motel chain's revenues during the next four quarters. Assume that advertisingduring each of the next four quarters is $50,000. (Hint: Try using advertising, lagged by one period, as an explanatory variable. See the Problem 60 for an explanation of a lagged variable. Also, use dummy variables for the quarters to account for possible seasonality.) b. Use simple exponential smoothing to make predic- tions for the motel chain's revenues during the next four quarters. Experiment with the smoothing constant. c. Use Holt's method to make forecasts for the motel chain's revenues during the next four quarters. Experiment with the smoothing constants. d. Use Winters' method to determine predictions for the motel chain's revenues during the next four quarters. Experiment with the smoothing constants. e. Which forecasts from parts a to d would you expect to be the most reliable? 20. The file P13_20.xlsx contains the monthly sales of iPod cases at an electronics store for a two-year period. Use the moving averages method, with spans of your choice, to forecast sales for the next six months. Does this method appear to track sales well? If not, what might be the reason? 26. The file P13_26.xlsx contains the monthly number of airline tickets sold by the CareFree Travel Agency. a. Create a time series chart of the data. Based on what you see, which of the exponential smoothing models do you think will provide the best forecast- ing model? Why? b. Use simple exponential smoothing to forecast these data, using a smoothing constant of 0.1. c. Repeat part b, but search for the smoothing con- stant that makes RMSE as small as possible. Does it make much of an improvement over the model in part b? 28. The file P13_28.xlsx contains monthly retail sales of U.S. liquor stores. a. Is seasonality present in these data? If so, charac- terize the seasonality pattern. b. Use Winters' method to forecast this series with smoothing constants a = B = 0.1 and y = 0.3. Does the forecast series seem to track the seasonal pattern well? What are your forecasts for the next 12 months?See Answer
  • Q6: Principles of Business Data Mining Project This group project offers you an opportunity to apply your data mining knowledge to real-life data and to mine managerially relevant insights. The objective of this assignment is for your team to implement the data mining process using real-life data that is of interest to you. Your project should be driven by a relevant and important question of business or social value. It is essential that the dataset you select have an output (also called target) variable. Note however, it is not enough to just have a target variable. Your dataset should also contain an adequate number of input variables which can help explain or predict the target variable. Data can be obtained from a publicly available source such as kaggle.com or through a real business (for which you need to have appropriate Non-Disclosure Agreements in place). You will need to set up an account to download datasets from kaggle which are usually in the csv format. Pay attention to the FIVE project deadlines: project proposal report, meeting to discuss project, project progress report, project presentation slides/presentation and final project report. All four reports/slides (one copy per team) -- the proposal, progress report, presentation slides and the final report -- should be submitted in Canvas. All late submissions will receive a zero. Each report should be a single Microsoft Word or pdf document with your group number and all group member names. Messy or hard-to-read reports will be penalized. You have to implement the following data mining techniques in your project: • • At least 2 data visualization techniques using Tableau to understand and draw conclusions on the data At least 2 prediction techniques to address your business or social question. I would recommend using a linear or logistic regression and classification trees since they are the most straightforward to translate to a relevant business question. Regardless of what data you will be mining, ensure that there is an appropriate match between the dataset you plan to mine and the data mining technique you plan to use and the business question. Deliverables: Project Proposal Report : ○ ○ Source of the dataset (e.g. kaggle.com) The proposal should address the following: A brief description of the data so you know what you are dealing with. You should include a list of all variables in the dataset. A short paragraph describing your objectives when mining this data. Explain what business problem/question, this project will address. What data visualization and prediction techniques you will use. Any pre-processing steps you think you need to take. 1 INSY5339 Dr. A.C. Sahoo Spring 2022 ○ Any initial results you expect or may have obtained. I strongly recommend that you run some of your data visualization techniques by the proposal due date and have a good view of what prediction techniques you will use. Project Progress Report: о ○ ○ At the beginning of this report, describe your business/social question and the data that support the addressing of this question. Data should be available and submitted as a separate file along with the proposal. The report should contain all the data visualization techniques (at least two) successfully completed and documented (it could change in the final report). Include your initial draft findings (it could change in the final report). Your report should contain at least one successful trial of the data prediction techniques (out of two techniques). You should report your draft findings (it could change in the final report). Project Presentation slides and presentation in the class ○ Your group will make a 15 minute presentation during class on December 1, 2021. Detailed guidelines will be presented prior to the final presentation. A copy of the presentation slides have to be submitted by the due date. ○ Every member should speak as part of the presentation. Professional quality presentation slides are required and should be updated based on feedback received. Final Project Report: ○ ○ This should be a professionally prepared report that addresses the following parts: cover page, executive summary, project motivation/background (business/social question), data description that supports addressing this question, data analysis using visualization and findings, your prediction models and findings, managerial or policy implications and conclusions, include all diagrams, graphs and tables to support your conclusions. Feel free to add any other sections if needed. What really matters is whether you successfully discovered useful knowledge from a dataset, and whether you presented it well to reader. Each report builds on the previous one. Feel free to reuse material in your earlier reports. Your Project Proposal Report submission should also contain the following completed table: How many observations in the dataset? How many binary/categorical variables? How many continuous variables? What is the outcome / target variable? 2 INSY5339 If binary or categorical: What percentage of the variables belong to each class. If continuous: What is the mean value of the target variable? Before doing any further processing, what would your prediction of the target variable be? 3 Dr. A.C. Sahoo Spring 2022See Answer
  • Q7: Data Mining Week 8 Instructions Address the following questions in complete sentences and submit the answers in a Word document. Problem 1: You are an operations analyst working for a major movie theater chain. Currently, the company requires all employees in the concession stand to use a technique known as suggestive selling (e.g., "Would you like Red Vines with your order today?"). Employees are not given any guidelines as to what suggestion to make, so they typically pick their favorite food or candy. The company has asked you to come up with a system to find items that a given customer is likely to buy in order to improve the effectiveness of the suggestive selling process and increase profits. Using market basket analysis, you have discovered several associations that you believe will lead to improved suggestive selling accuracy. You are considering two main options for deployment: 1. Work with the point-of-sale software vendor to implement the suggestive selling algorithm into the software used by cashiers. As cashiers enter the customer's order, the algorithm will determine what additional item the customer is likely to buy, based on the items currently ordered. The item will be displayed to the cashier who can then choose to add the item to the current order or dismiss the prompt, depending on the customer's response. Naturally, the software vendor is demanding a hefty fee and promising a 6-12-month timeline for delivery of the software update. Changes to the algorithm down the road will require an additional fee and waiting period. 2. Choose a small number (e.g., 5-10) of the most effective rules and incorporate them into the training protocol. One of these rules would be a catch-all (i.e., "If no other rules apply, suggest Red Vines"). You estimate that it will take 2-4 weeks to decide on rules, develop training collateral (posters for the employee breakroom, reminder cards to stick on each cash register, etc.), and deliver training. Because this is a high turnover industry and training is constantly being delivered, the additional costs of delivering this training are negligible. In 100-250 words, describe what sort of analysis you would need to do to make an informed decision about what plan to implement. What data would you need to acquire? How would you acquire that data? How would you measure the success of the deployment? How would you determine when the model needs to be updated? Problem 2: You have decided to implement one of the plans from the previous exercise. The company has asked you to write a 250-word e-mail message that will be sent out to all managers in the company announcing the upcoming change. While managers understand movie theater operations, they do not have a background in data mining. Your e-mail should clearly explain all of the following in layman's terms. You may make up figures related to costs, revenue increases, etc., as needed, to support your e-mail communication. 1. Type of model being implemented 2. Purpose for implementing the model 3. Anticipated results 4. Next steps and timeline for implementation Problem 3: Read the case study "Championing of an LTV Model at LTC," located in topic Resources, and answer the following questions in complete sentences. 1. What was the business problem the authors were trying to solve? 2. What type of modeling activities did the authors use? (description, prediction, classification, etc.) 3. How did the authors evaluate their model? 4. What stages of CRISP-DM are not represented in this case study? 5. Based on this article, why do you think it is important that implementers of data mining models possess strong interpersonal skills? 6. Consider the "Data Science Code of Professional Conduct" and evaluate two ethical issues related to data mining and the responsible stewardship of personal information. https://www.researchgate.net/publication/220520009 Championing of_an_LTV_model_ at LTCSee Answer
  • Q8: 2. In PC Tech's product mix problem, assume there is another PC model, the VXP, that the company can produce in addition to Basics and XPs. Each VXP requires eight hours for assembling, three hours for testing, $275 for component parts, and sells for $560. At most 50 VXPs can be sold. a. Modify the spreadsheet model to include this new product, and use Solver to find the optimal product mix. 4. Again continuing Problem 2, suppose that you want to force the optimal solution to be integers. Do this in Solver by adding a new constraint. Select the deci- sion variable cells for the left side of the constraint, and in the middle dropdown list, select the "int" op- tion. How does the optimal integer solution compare to the optimal noninteger solution in Problem 2? Are the decision variable cell values rounded versions of those in Problem 2? Is the objective value more or less than in Problem 2? 26. A furniture company manufactures desks and chairs. Each desk uses four units of wood, and each chair uses three units of wood. A desk contributes $250 to profit, and a chair contributes $145. Marketing restrictions require that the number of chairs produced be at least four times the number of desks produced. There are 2000 units of wood available. a. Use Solver to maximize the company's profit. 46. During each four-hour period, the Smalltown police force requires the following number of on-duty police officers: four from midnight to 4 A.M.; four from 4 A.M. to 8 A.M.; seven from 8 A.M. to noon; seven from noon to 4 P.M.; eight from 4 P.M. to 8 P.M.; and ten from 8 P.M. to midnight. Each police officer works two consecutive four-hour shifts. a. Determine how to minimize the number of police of- ficers needed to meet Smalltown's dailyrequirements. 66. United Steel manufactures two types of steel at three different steel mills. During a given month, each steel mill has 240 hours of blast furnace time available. Because of differences in the furnaces at each mill, the time and cost to produce a ton of steel differ for each mill, as listed in the file P04_66.xlsx. Each month, the company must manufacture at least 700 tons of steel 1 and 600 tons of steel 2. Determine how United Steel can minimize the cost of manufacturing the desired steel.See Answer
  • Q9: 2. In PC Tech's product mix problem, assume there is another PC model, the VXP, that the company can produce in addition to Basics and XPs. Each VXP requires eight hours for assembling, three hours for testing, $275 for component parts, and sells for $560. At most 50 VXPs can be sold. a. Modify the spreadsheet model to include this new product, and use Solver to find the optimal product mix. 4. Again continuing Problem 2, suppose that you want to force the optimal solution to be integers. Do this in Solver by adding a new constraint. Select the deci- sion variable cells for the left side of the constraint, and in the middle dropdown list, select the "int" op- tion. How does the optimal integer solution compare to the optimal noninteger solution in Problem 2? Are the decision variable cell values rounded versions of those in Problem 2? Is the objective value more or less than in Problem 2? 26. A furniture company manufactures desks and chairs. Each desk uses four units of wood, and each chair uses three units of wood. A desk contributes $250 to profit, and a chair contributes $145. Marketing restrictions require that the number of chairs produced be at least four times the number of desks produced. There are 2000 units of wood available. a. Use Solver to maximize the company's profit. 46. During each four-hour period, the Smalltown police force requires the following number of on-duty police officers: four from midnight to 4 A.M.; four from 4 A.M. to 8 A.M.; seven from 8 A.M. to noon; seven from noon to 4 P.M.; eight from 4 P.M. to 8 P.M.; and ten from 8 P.M. to midnight. Each police officer works two consecutive four-hour shifts. a. Determine how to minimize the number of police of- ficers needed to meet Smalltown's dailyrequirements. 66. United Steel manufactures two types of steel at three different steel mills. During a given month, each steel mill has 240 hours of blast furnace time available. Because of differences in the furnaces at each mill, the time and cost to produce a ton of steel differ for each mill, as listed in the file P04_66.xlsx. Each month, the company must manufacture at least 700 tons of steel 1 and 600 tons of steel 2. Determine how United Steel can minimize the cost of manufacturing the desired steel.See Answer
  • Q10: Each student is required to submit (1) the Orange workflow file or Python/R script files (s)he created using the format HW1PID.OWS (.py and .r respectively for Python/R) where PID is your PID number, and (2) a PDF file with answers using the format HW3PID.pdf All files listed below were posted on Canvas sub-folder Homeworks of Data sets folder. Homework III - Questions The purpose of this homework is to use textual analysis of WSJ news in predicting financial market outcomes. In particular we will rely on a data set measuring the state of the economy via textual analysis of business news. From the full text content of 800,000 Wall Street Journal articles for 1984-2017, Bybee et al. estimate a model that summarizes business news as easily interpretable topical themes and quantifies the proportion of news attention allocated to each theme at each point in time. These news attention estimates are inputs into the models we want to estimate. The data source is described in http://structureofnews.com/. The data was standardized and prepared for this assignment. Please use the data file ML_MBA_UNC_Processed.xlxs from the HW3 material folder (sub-folder of Assignments folder) on CANVAS (it is also posted in the Data Sets folder - Homework sub-folder) Click on the News Taxomony tab of the aforementioned website and you will find a taxon- omy of news themes in The Wall Street Journal. Estimated with hierarchical agglomerative clustering, the dendrogram illustrates how 180 topics cluster into an intuitive hierarchy of increasingly broad metatopics. The list of topics is reproduced below - further details appear on the website: Natural disasters, Internet, Soft drinks, Mobile devices, Profits, M&A, Changes, Police / crime, Research, Executive pay, Mid- size cities, Scenario analysis, Economic ideology, Middle east, Savings & loans, IPOs, Restraint, Electronics, Record high, Connecticut, Steel, Bond yields, Small business, Cable, Fast food, Disease, Activists, Competition, Music industry, Short sales, Nonperforming loans, Key role, News conference, US defense, Political contributions, Revised estimate, Economic growth, Justice Department, Credit ratings, Broadcasting, Problems, Announce plan, Federal Reserve, Job cuts, Chemicals / paper, Regulation, Environment, Small caps, Unions, C-suite, Control stakes, Mutual funds, Venture capital, European sovereign debt, Mining, Company spokesperson, Private / public sec- tor, Pharma, Schools, Russia, Programs / initiatives, Health insurance, Drexel, Trade agreements, Treasury bonds, Challenges, People familiar, Sales call, Publishing, Financial crisis, Aerospace / defense, Recession, Latin America, Cultural life, SEC, Earnings losses, Phone companies, Computers, Marketing, Japan, Nuclear / North Korea, NY politics, Tobacco, Product prices, Biology / chemistry / physics, Movie industry, Automotive, Machinery, Bankruptcy, Arts, International exchanges, Accounting, Space program, Immigration, Small changes, Small possibility, Agreement reached, Oil drilling, Rail / trucking / shipping, Indictments, Positive sentiment, Canada / South Africa, Airlines, California, Corporate governance, China, Investment banking, Spring/summer, Software, Pensions, Humor / language, Systems, Clintons, Major concerns, Mid-level executives, US Senate, Agriculture, Bank loans, Takeovers, State politics, Real estate, Futures / indices, Southeast Asia, Optimism, Corrections / amplifications, Government budgets, Exchanges / composites, Currencies / metals, Mortgages, Financial reports, Germany, Rental properties, Committees, Subsidiaries, Management changes, Share payouts, France / Italy, Acquired investment banks, Credit cards, Bear / bull market, Earnings forecasts, Terrorism, Watchdogs, Oil market, Couriers, Commodities, Utilities, Foods / consumer goods, Convertible / preferred, Macroeconomic data, Courts, Safety admin- istrations, Reagan, Bush / Obama / Trump, Fees, Gender issues, Trading activity, Microchips, Insurance, Earnings, Luxury / beverages, Iraq, National security, Buffett, Taxes, Options / VIX, Casinos, Elections, Private equity / hedge funds, Negotiations, European politics, Size, NASD, Mexico, Retail, Long / short term, Wide range, Lawsuits, UK, Revenue growth 1.a [4 points] For the exercise below select ALL features EXCEPT FUTSP500, APPL, SGNAPPL, FUTAPPL, SGNFUTAPPL, SGNSP500, DATE and use as target: SGNFUTSP500 Your first task is to predict the direction of the market (S&P 500 index) over the next month, i.e. SGNFUTSP500 up or down - which is a classification prediction problem, using the importance of news topics. Use the logistic regression and neural network widgets of the Orange software to compute the following models: logistic regression with LASSO regularization with C = 0.007 logistic regression with LASSO regularization with C = 0.80 • neural network with 100 hidden layers, ReLu activation function, SGD optimization, and max 200 iterations and regularization a = = 5 You should get something along the following lines with 10-fold cross-validation: AUC CA ● Models Logistic LASSO C = 0.80 Neural Net Logistic LASSO C = 0.007 0.583 0.649 0.566 0.619 0.500 0.371 Explain why the Logistic regression with LASSO C = 0.007 does so poorly (hint: look at the coefficients of the model). When you look at the ROC curve, explain the curve for the LASSO C = 0.007 model. 2 1.b [4 points] List the top ten features which have the most negative impact on next month's market direction and the top ten with the most positive impact. 2.a [4 points] We are turning now to a linear regression instead of classification problem, predicting actual market returns (continuous) rather than direction (binary). For the exer- cise below select ALL features EXCEPT SGNFUTSP500, APPL, SGNAPPL, FUTAPPL, SGNFUTAPPL, SGNSP500, DATE and use as target: FUTSP500 Use the linear regression and neural net widgets of Orange software to compute the follow- ing models: regression with Elastic Net with a = 0.007 for the LASSO/Ridge regularization and weight 0.80 on the l₂ regularization. Note that the logistic regression and linear regression widgets use a different way of writing the penalty function (although there is a mapping between C and a via something called the Lagrangian multiplier not covered in the course). neural network with 100 hidden layers, ReLu activation function, SGD optimization, and max 200 iterations and regularization a = = 5 The Test and Score widget now records MSE, RMSE, MAE, and R2. What do you learn from the R2 results? 2.b [4 points] In the previous case we were trying to predict the return next month with current news. Now, we will try to explain current returns with current news. For the exercise below select ALL features EXCEPT FUTSP500, SGNFUTSP500, APPL, SGNAPPL, FUTAPPL, SGNFUTAPPL, SGNSP500, DATE and use as target: SP500 Use the linear regression widget of Orange software to compute the following models: regression with Elastic Net with a = 0.007 for the LASSO/Ridge regularization and weight 0.80 on the l₂ regularization. Recall that the logistic regression and linear re- gression widgets use a different way of writing the penalty function (although there is a mapping between the role of C and a via something called the Lagrangian mul- tiplier not covered in the course). 3 • neural network with 100 hidden layers, ReLu activation function, SGD optimization, and max 200 iterations and regularization a = 5 The Test and Score widget now records MSE, RMSE, MAE, and R2. What do you learn from the R2 results? 2.c [2 points] List the top ten features which have the most negative impact on next month's market returns and the top ten with the most positive impact. How is it different from the results in 1.b? 2.d [2 points] Explain the difference between your answers in 1.a, 2.a and 2.b. 4See Answer
  • Q11:2. Use OPTICS algorithm to output the reachability distance and the cluster ordering for the dataset provided, starting from Instance 1. Use the following parameters for discovering the cluster ordering: minPts =2 and epsilon =2. Use epsilonprime =1.2 to generate clusters from the cluster ordering and their reachability distance. Don't forget to record the core distance of a data point if it has a dense neighborhood. You don't need to include the core distance in your result but you may need to use them in generating clusters. (45 pts) Instance 1: Instance 2: Instance 3: Instance 16: Dataset visualization Below are the first few lines of the calculation. You need to complete the remaining lines and generate clusters based on the given epsilonprime value: Instance (X,Y) (1, 1) (0, 1) (1, 0) (5,9) 617 Reachability Distance Undefined (or infinity) 1.0 1.0 UndefinedSee Answer
  • Q12:1. If Epsilon is 2 and minpoint is 2 (including the centroid itself), what are the clusters that DBScan would discover with the following 8 examples: A1=(2,10), A2=(2,5), A3=(8,4), A4=(5,8), A5=(7,5), A6=(6,4), A7=(1,2), A8=(4,9). Use the Euclidean distance. Draw the 10 by 10 space and illustrate the discovered clusters. What if Epsilon is increased to sqrt(10)? (30 pts)See Answer
  • Q13:2the Task - Correlation Rule Extraction and Outlier Detection Prices Question 1 Suppose we have ten product codes (from 0 to 9) and the following table of transactions Identifier Transaction 1 2 3 4 5 6 7 8 9 10 Products 0, 1, 3, 4 1, 2, 3 1, 2, 4, 5 1, 3, 4, 5 2, 3, 4, 5 2, 4, 5 3,A 1, 2, 3 1, 4, 5 B,4 where, the penultimate and last digit of your registration number divided by2 and rounded to the nearest unit. That is, if your registration number ends in 12, then you will set = 1/2= 0.5 that is1 and =2/2= 1. But if the penultimate digit of the number your registry is either 5 or 6, you will set = 4. Accordingly, if the last digit of the number your register is either 3 or 4, you will set = 2. Please answer the following questions, showing your calculations: 1. Compute the support of the sets {4,5} and {3, 5} 2. The support and trust of the rules {4, 5} → {3} and {3, 5} → {2} 3. Run the Apriori Algorithm with the method Fk-1X F k-1 for support threshold equal to 2 and report which frequent sets you find at each stepSee Answer
  • Q14:This tutorial will guide you how to do homework in this course. 1. Goto https://www.kaggle.com/c/titanic and follow walkthrough as https://www.kaggle.com/alexisbcook/titanic-tutorial B 2. Submit your result to Kaggle challenge. 3. Post jupyter notebook to your homepage as blog post. A good example of blog post is. https://jalammar.github.io/visual-interactive-guide-basics-neural-networks/ 4. Submit your homepage link and screenshot pdf in the canvas. 5. Doing 1-4 will give you 8 points. To get additional 2 points, create a section as "Contribution" and try to improve the performance. I expect one or two paragraph minimum (the longer the better). Show the original score and improved score.See Answer
  • Q15:1. We will use Flower classification dataset a. https://www.kaggle.com/competitions/tpu-getting-started 2. Your goal is improving the average accuracy of classification. a. You SHOULD use google collab as the main computing. (Using Kaggle is okay) b. You SHOULD create a github reposit for the source code i. Put a readme file for execution c. You SHOULD explain your source code in the BLOG. d. Try experimenting with various hyperparameters i. Network topology 1. Number of neurons per layer (for example, 100 x 200 x 100, 200 x 300 x 100...) 2. number of layers (For example, 2 vs 3 vs 4 ... ) 3. shape of conv2d ii. While doing experiments, make sure you record your performance such that you can create a bar chart of the performance iii. An additional graph idea might be a training time comparison Do some research on ideas for improving this. iv. e. You can refer to the code or tutorial internet. But the main question you have to answer is what improvement you made over the existing reference. i. Make sure it is very clear which lines of code is yours or not. When you copy the source code, add a reference. 3. Documentation is the half of your work. Write a good blog post for your work and step-by-step how to guide. a. A good example is https://jalammar.github.io/visual-interactive-guide-basics-neural-networks/ 4. Add a reference a. You add a citation number in the contents and put the reference in the separate reference sectionSee Answer
  • Q16:3. Use F-measure and the Pairwise measures (TP, FN, FP, TN) to measure the agreement between a clustering result (C1, C2, C3) and the ground truth partitions (T1, T2, T3) as shown below. Show details of your calculation. (25 pts) Ground Truth T, TT, Cluster C, CC3See Answer
  • Q17: 2. Use OPTICS algorithm to output the reachability distance and the cluster ordering for the dataset provided, starting from Instance 1. Use the following parameters for discovering the cluster ordering: minPts =2 and epsilon =2. Use epsilonprime =1.2 to generate clusters from the cluster ordering and their reachability distance. Don't forget to record the core distance of a data point if it has a dense neighborhood. You don't need to include the core distance in your result but you may need to use them in generating clusters. (45 pts) 2 16 14 12 10 015 05 20 06 04 021 026 016 027 022 025 019 023 09 024 07 018 08 011 070 030 029 028 012 013 014 2 01 0 03 0 2 4 8 10 12 14 16 017 Dataset visualization Below are the first few lines of the calculation. You need to complete the remaining lines and generate clusters based on the given epsilonprime value: Instance (X,Y) Reachability Distance Instance 1: (1,1) Undefined(or infinity) Instance 2: (0, 1) 1.0 Instance 3: (1, 0) 1.0 Instance 16: (5,9) Undefined Instance 13: (9,2) Undefined Instance 12: (8,2) 1See Answer
  • Q18:Assignment #3: DBSCAN, OPTICS, and Clustering Evaluation 1. If Epsilon is 2 and minpoint is 2 (including the centroid itself), what are the clusters that DBScan would discover with the following 8 examples: A1=(2,10), A2=(2,5), A3=(8,4), A4=(5,8), A5=(7,5), A6=(6,4), A7=(1,2), A8=(4,9). Use the Euclidean distance. Draw the 10 by 10 space and illustrate the discovered clusters. What if Epsilon is increased to sqrt(10)? (30 pts)See Answer
  • Q19:Discussion - Data Mining, Text Mining, and Sentiment Analysis Explain the relationship between data mining, text mining, and sentiment analysis. Provide situations where you would use each of the three techniques. Respond to the following in a minimum of 230 words:See Answer
  • Q20:Please create a K-means Clustering and Hierarchical Clustering with the line of code provided. The line of code should include a merger of the excel files. The excel files will also be provided See Answer

TutorBin Testimonials

I found TutorBin Data Mining homework help when I was struggling with complex concepts. Experts provided step-wise explanations and examples to help me understand concepts clearly.

Rick Jordon

5

TutorBin experts resolve your doubts without making you wait for long. Their experts are responsive & available 24/7 whenever you need Data Mining subject guidance.

Andrea Jacobs

5

I trust TutorBin for assisting me in completing Data Mining assignments with quality and 100% accuracy. Experts are polite, listen to my problems, and have extensive experience in their domain.

Lilian King

5

I got my Data Mining homework done on time. My assignment is proofread and edited by professionals. Got zero plagiarism as experts developed my assignment from scratch. Feel relieved and super excited.

Joey Dip

5

TutorBin helping students around the globe

TutorBin believes that distance should never be a barrier to learning. Over 500000+ orders and 100000+ happy customers explain TutorBin has become the name that keeps learning fun in the UK, USA, Canada, Australia, Singapore, and UAE.