tutorbin

advanced statistics homework help

Boost your journey with 24/7 access to skilled experts, offering unmatched advanced statistics homework help

tutorbin

Trusted by 1.1 M+ Happy Students

WhatsApp Support

Get Instant
Online Homework Help
via WhatsApp

Get instant homework help from top tutors—just a WhatsApp message away. 24/7 hw help support for all your academic needs!

A
S
M
R
★★★★★
2M+ students trust TutorBin
Your WhatsApp Number
phone
or
⚡ Instant reply
🔒 100% private
👨‍🏫 Top tutors
🌍 All subjects
*Get instant homework help from top tutors—just a WhatsApp message away. 24/7 support for all your academic needs!
2M+ Students Helped24/7 Live SupportExpert TutorsAll Subjects CoveredInstant Response100% ConfidentialTop Rated ServiceMoney-back Guarantee2M+ Students Helped24/7 Live SupportExpert TutorsAll Subjects CoveredInstant Response100% ConfidentialTop Rated ServiceMoney-back Guarantee

Recently Asked advanced statistics Questions

Expert help when you need it
  • Q1:Assignment-3 MSCA31010: Linear & Non-Linear Models (Acknowledgement: Special thanks to Francisco Azeredo & Ming_Long Lam for their content creation support) For this homework, you can use either R or Python, use Word docs, R-markdowns, Jupyter notebooks, or html’s for submission. Train a binary logistic regression model on the claim_history.csv. Your model will predict the likelihood of filing more than one claim in one unit of exposure. You will first calculate the Frequency variable by dividing the CLM_COUNT by EXPOSURE. Next, you will create a binary target variable that determines if the Frequency is strictly greater than one (i.e., the Event). You will use MSTATUS, CAR_TYPE, REVOKED, and URBANICITY as the categorical predictors, and CAR_AGE, MVR_PTS, TIF, and TRAVTIME as the interval predictors. Your goal is to train a model that has just the right set of predictors. The standard libraries for R or Python are allowed. You need to drop all missing values (i.e., NaN) of all the predictors and the target variable before training your model. (15 points) Before you train the model, we want to explore the predictors. For each predictor, generate a line chart that shows the odds of the Event by the predictor’s unique values. The predictor’s unique values are displayed in ascending lexical order. (20 points) Enter the predictors into your model using Forward Selection. The Entry Threshold is 0.05. Please provide a detailed report of the Forward Selection. However, you do not need to show steps such as in the previous question. The report should include (1) the predictor entered, (2) the log-likelihood value, (3) the Deviance Chi-squares statistic, (4) the Deviance Degree of Freedom, and (5) the Chi-square significance. (10 points). Which predictors does your final model contain? (10 points). Please show a table of the complete set of parameters of your final model. Please also include the exponentiated estimates (i.e., apply the exp() function on the parameter estimates). 2. You will visually assess your final model in Question 1. Please color-code the markers according to the Exposure value. Also, please briefly comment on the graphs. (10 points). Please plot the predicted Event probability versus the observed Frequency. (10 points). Please plot the Deviance residuals versus the observed Frequency. 3. (15 Points) You will calculate the Accuracy metric to assess your final model in Question 3. If the predicted Event probability of an observation is greater than or equal to 0.25, then you will classify that observation as the Event (i.e., filing more than one claim per unit exposure). An observation is correctly classified if the predicted target value equals the observed target value. The Accuracy metric is the proportion of observations that are correctly classified. Bonus: (20 Points) For questions 1B, 1C and 1D apply recursive feature elimination (RFE) instead of Forward Selection (see: https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html )See Answer
  • Q2:Case studies should be formatted according to APA guidelines for an executive summary. ul. Case Study Alumni donations are an important source of revenue for colleges and universities. If administrators could determine the factors that could lead to increases in the percentage of alumni who make a donation, they might be able to implement policies that could lead to increased revenues. Research shows that students who are more satisfied with their contact with teachers are more likely to graduate. As a result, one might suspect that smaller class sizes and lower student/faculty ratios might lead to a higher percentage of satisfied graduates, which in turn might lead to increases in the percentage of alumni who make a donation. The following table shows data for 48 national universities. The Graduation Rate column is the percentage of students who initially enrolled at the university and graduated. The % of Classes Under 20 column shows the percentages of classes with fewer than 20 students that are offered. The Student/Faculty Ratio column is the number of students enrolled divided by the total number of faculty. Finally, the Alumni Giving Rate column is the percentage of alumni who made a donation to the university. • AlumniGiving.xlsx Directions Respond in detail to each question. Review the rubric prior to responding. You can use R or JMP, and refer to the chapter readings for additional regression model suggestions. Prepare a managerial report as described below. In the Excel data, add in four additional rows for Add in Park University, Donnelly College, <Instructor Last Name> University, and University of <Your Last Name> 1. Use methods of descriptive statistics to summarize the data. 2. Develop an estimated simple linear regression model that can be used to predict the alumni giving rate, given the graduation rate. Describe and analyze your findings. 3. Develop an estimated multiple linear regression model that could be used to predict the alumni giving rate using Graduation Rate, % of Classes Under 20, and Student/ Faculty Ratio as independent variables. Discuss your findings. 4. Based on the results in parts (2) and (3), do you believe another regression model may be more appropriate? Estimate this model, and discuss your results. 5. What conclusions and recommendations can you derive from your analysis? What universities are achieving a substantially higher alumni giving rate than would be expected, given their Graduation Rate, % of Classes Under 20, and Student/Faculty Ratio? What universities are achieving a substantially lower alumni giving rate than would be expected, given their Graduation Rate, % of Classes Under 20, and Student/ Faculty Ratio? What other independent variables could be included in the model? Remember to include screenshots of the models discussed in your paper. Assignment Instructions ||documents required word and excel in APA format.See Answer
  • Q3:Part 2: You are a business analytics analyst, and your director has given you the task of analyzing automobile data. Using the attached "Automobiles" dataset, perform a K-Means cluster analysis on these variables. It is recommended that you perform a cluster analysis for 2, 3, 4, and 5 clusters and select the optimal model based on the results. In a 300-word summary, provide the requested screenshots and address the following: 1. Explain your approach to the problem. 2. State the optimal number of clusters and the cluster sizes, and discuss the most logical cluster characteristics. Include discussion of how you would explain the attributes of each cluster group and justify your explanation. 3. Provide at least three scatterplots of selected input variables of your choice that help explain the cluster categories. 4. Provide a line chart of the silhouette coefficients for each cluster and the overall silhouette coefficient for the model. 5. Export the cluster assignment for each record to an Excel file. Name the file "Clusters.xlsx." Provide overall conclusion based upon the results of the K-Means analysis. Note that you are required to submit the completed KNIME *.knwf file to your instructor. Specifically, export your KNIME model to a KNIME workflow file. To perform this task in KNIME, ensure that your KNIME model is active (i.e., displayed). Then, go to File - >Export KNIME Workflow. In the "Destination workflow file name (.knwf)" area, browse to a specific location on your computer. Click "Save" and then click "Finish."See Answer
  • Q4:Problem 2: (50 pts) You are given 204 observations from a travel survey conducted in the Seattle Metropolitan area. The purpose of the survey was to study the number of times (per week) commuters changed their departure time on their work-to-home trip to avoid traffic congestion. The data are non-negative integers with the mean approximately equal to the variance. Your task is to estimate the appropriate count-data model. • Following the forward stepwise process, find your best fit model specification. Some things to consider as you fit your model: - Is there acceptable correlation among explanatory variables? Do the signs of the coefficients make sense? You will need to create indicators for some of the variables. - Is the model you have chosen to use appropriate? Double-check after you have arrived at your best fit specifications. Explain your process of determining if the count-data model you are using is appropriate or not be specific. • After you have arrived at your best fit model specifications, provide a discussion of the logical process that led you to the selection of your final specification. • Present the descriptive statistics of the variables in your final model specification. You do not have to categorize variables by category, but are welcome to do so. Points deducted for incorrect model presentation. • Present your model as shown in the document on Canvas. You do not have to categorize variables by category, but are welcome to do so. Points deducted for incorrect model presentation. • Provide a discussion for each of the variables in your final model specification, including their quantitative effect on your dependent variable. What are some plausible reasons for the significance of the variable and its effect on changing the number of times (per week) a commuter changed their departure time? You are welcome to use your intuition or find sources that confirm/validate your results. • Based on the distribution of your dependent variable, would a truncated model, cen- sored model, or zero-inflated model be appropriate? Explain and be specific. Definitions of variables are given on the following page.See Answer
  • Q5:Problem 1: (50 pts) You are given injury severity data for 2,273 single-vehicle motorcycle crashes in Indiana. There are four possible severity outcomes: no injury (property damage only and possible injury), non-incapacitating injury, incapacitating injury, and fatality. You want to know what the likelihood of an individual being involved in a crash is based on the available crash data characteristics. Your task is to estimate an ordered probability model of motorcyclists' injury severity. • Following a forward stepwise process, find your best fit model specification. Some things to consider as you fit your model: Make sure the dependent variable is ordered and follows the format required for model estimation. - Do the signs of the coefficients make sense? - Is there acceptable correlation among explanatory variables? You must create some indicators that I did not cover in the tutorial documents. - Do you need to correct for heteroskedasticity? • After you have arrived at your best fit model specifications, provide a detailed discus- sion of the logical process that led you to the selection of your final specification. Present the descriptive statistics of the variables included in your final model specifica- tions as shown in the document on Canvas. You do not have to categorize the variables by category, but are welcome to do so. Are there any statistics worth highlighting? Points will be deducted for incorrect descriptive statistics presentation. • Present your model as shown in the document on Canvas. You do not have to categorize the variables by category, but are welcome to do so. Points will be deducted for incorrect model presentation. • Provide a discussion for each of the variables in your final model specifica- tion. Based on your model, what are the effects of the variables on fatal/injury crash probability? What are some plausible reasons for the significance of the variable and its effect on injury severity outcome? You are welcome to use your intuition, but also encouraged to find sources that confirm/validate your results. • Summarize your findings and provide some potential solutions for at least three of the variables in your final model specification. Use the countermeasure selection resources to determine your proposed solutions, explain why a countermeasure may be effective, and what the anticipated increase in safety would be should the countermeasure be implemented. • Prepare your deliverable as a mini-report based on what is being asked in the bullet points above. Definitions of available variables are given on the following page.See Answer
  • Q6:Problem 2. (15 points) Suppose Y~ Nn (µ, In). Let X € Rnxp be a fixed matrix with full column rank. In this exercise we consider estimators of the form û = Xß for some estimator 3. (i) Let μLS XLS be the corresponding estimator of the mean of Y based on the least squares estimator BLS = (XX)-¹XTY. Find the risk of this estimator directly in terms of Co= X(XTX)-¹XT, μ, n, and p only. (ii) Suppose μ = XB* for some ß* ERP. Show that R(μ, ûLS) = p. = = (iii) Consider the ridge estimator pridge = (X¹X+8Ip)-¹X¹Y and the corresponding matrix Cg = X(XTX + 8Ip)-¹X, where 820. Show that R(µ, pridge) = tr(C3) + || (In-Cs)μ||². (iv) Suppose that XTX = Ip and that μ = XB* for some ß*. Show that for every 3* 0 there exists 8 >0 such that R(XB*, pridge) < R(XB*, μLS).See Answer
  • Q7:/nInstructions Case study should be formatted according to APA guidelines for an executive summary. Respond in detail to each question. Review the rubric prior to responding. Prepare a managerial report as described below. Develop a model that will allow Applecore to maximize the number of customers reached for a budget of $10,000 for one week of promotion. Solve the model. What is the maximum number of customers reached for the $10,000 budget? Perform a sensitivity analysis on the budget for values from $5,000 to $48,000 in increments of $4,000. Construct a graph of percentage reach versus budget. Is the additional increase in percentage reach monotonically decreasing as the budget allocation increases? Why or why not? What is your recommended budget? Explain. Word limit for report - minimum 400 words/n ul. Case Study Applecore Children's Clothing is a retailer that sells high-end clothes for toddlers (ages 1 to 3), primarily in shopping malls. Applecore also has a successful Internet-based sales division. Recently Dave Walker, vice-president of the e-commerce division, has been given the directive to expand the company's Internet sales. He commissioned a major study on the effectiveness of Internet ads placed on news web sites. The results were favorable: Current patrons who purchased via the Internet and saw the ads on news web sites spent more, on average, than did comparable Internet customers who did not see the ads. With this new information on Internet ads, Walker continued to investigate how new Internet customers could most effectively be reached. One of these ideas involved strategically purchasing ads on news web sites prior to and during the holiday season. To determine which news sites might be the most effective for ads, Walker conducted a follow-up study. An e-mail questionnaire was administered to a sample of 1,200 current Internet customers to ascertain which of 30 news sites they regularly visit. The idea is that web sites with high proportions of current customer visits would be viable sources of future customers for Applecore products. Walker would like to ascertain which news sites should be selected for ads. The problem is complicated because Walker does not want to count multiple exposures. So, if a respondent visits multiple sites with Applecore ads or visits a given site multiple times, that respondent should be counted as reached but not more than once. In other words, a customer is considered reached if he or she has visited at least one web site with an Applecore ad. Data from the customer e-mail survey have begun to trickle in. Walker wants to develop a prototype model based on the current survey results. So far, 53 surveys have been returned. To keep the prototype model manageable, Walker wants to proceed with model development using the data from the 53 returned surveys and using only the first 10 news sites in the questionnaire. The costs of ads per week for the 10 web sites are given in the following table, and the budget is $10,000 per week. For each of the 53 responses received, the 10 web sites visited regularly are shown below. For a given customer-web site pair, a one indicates that the customer regularly visits that web site and a zero indicates that the customer does not regularly visit that site. • Applecore (Excel File) Directions Respond in detail to each question. Review the rubric prior to responding. Prepare a managerial report as described below. 1. Develop a model that will allow Applecore to maximize the number of customers reached for a budget of $10,000 for one week of promotion. 2. Solve the model. What is the maximum number of customers reached for the $10,000 budget? 3. Perform a sensitivity analysis on the budget for values from $5,000 to $48,000 in increments of $4,000. Construct a graph of percentage reach versus budget. Is the additional increase in percentage reach monotonically decreasing as the budget allocation increases? Why or why not? What is your recommended budget? Explain. Knowledge and Skills Case Study Rubric Criteria Development with use of adequate support (text & outside sources) view longer description Research and analysis view longer description Interpretation and critical analysis view longer description APA and technical compliance view longer description Ratings 10 to >6.9 pts Meets Expectations Student included a minimum of 2 full pages of written content supported with 2 or more academic sources of research offering a detailed reflection and literature review on required topics for this paper. 30 to >20.9 pts Meets Expectations The response skillfully interprets, analyzes, and applies knowledge into skills in the case study in an insightful, original manner. The write up is edited to summarize the answer via a provided analytic framework and integration with required exercise elements. 35 to >25.9 pts Meets Expectations The paper contains in-depth scholarly detail offered conclusions detailing how knowledge and skills learned from required topics in this paper support continued professional and academic growth in aspects of data analysis and data analytics (when required, also includes visualizations in a process flow diagram in support of these conclusions) 5 to >3.9 pts Meets Expectations Response is clearly written with strong style and conforms to conventions of standard written English. Proper APA format with appropriate use of citations, quotations and other areas. No grammatical errors. 6.9 to >2.9 pts Developing Student included a minimum of 2 full pages of written content supported with some academic sources of research offering a detailed reflection and literature review on required topics for this paper. 20.9 to >10.9 pts Developing The response sufficiently interprets and analyzes the case study in adequate ways, though may lack originality and insight. Answers are longer or shorter because they are not edited as clearly and not reviewed as carefully. Analysis is clear, but slight improvement is needed. 25.9 to 14.9 pts Developing The paper contains some scholarly detail offered conclusions. Improvement is needed in detailing how knowledge and skills learned from required topics in this paper support continued professional and academic growth. There is some aspects of data analysis and data analytics (when required, also includes visualizations in a process flow diagram in support of these conclusions) 3.9 to 1.9 pts Developing Response is adequately written and conforms to conventions of standard written English. Mostly in APA format with appropriate use of citations, quotations and other areas. Minor or no grammatical errors. 2.9 to >0 pts No Marks Student did not include a minimum of 2 full pages of written content supported with no academic sources of research offering a detailed reflection and literature review on required topics for this paper. 10.9 to >0 pts No Marks The response interprets and analyzes the case (or research) in a predictable manner, summarizing ideas, rather than suggesting insights. There are some areas lacking. Analysis is unclear and structure is need of improvement. 14.9 to >0 pts No Marks The paper contains minimal scholarly detail offered conclusions. Somewhat meets the detailing requirements of application of knowledge and skills learned from required topics in this paper support continued professional and academic growth. There is minimal aspects of data analysis and data analytics (when required, also includes visualizations in a process flow diagram in support of these conclusions) 1.9 to >0 pts No Marks Response contains multiple errors not conforming to conventions of standard written English. There a lacking of APA format with appropriate use of citations, quotations and other areas. Grammatical errors. Pts / 10 pts / 30 pts / 35 pts / 5 ptsSee Answer
  • Q8: School of Criminology and Criminal Justice CRJ 511 Applied Data Analysis in Criminal Justice Application Assignment #4 For this assignment you will be using two articles. You are going to use Gray (2010) Actions for Part 1 and Dearing et al. (2005) Actions for Part 2. Again, as a reminder. Use the Gray (2010) article for questions 1-3 that are all related to chi-square testing. Use the Dearing et al. (2005) article for questions 4-9 that are related to correlations. If you have trouble accessing these links, let me know. I can email them to you directly! Unfortunately, these articles contain numerous chi-square test and correlation tables which could cause some confusion. To help, please keep the following in mind. Gray (2010)- I want you to answer the questions using Table III which is located on page 547. Here is a screenshot Download screenshotof the table. Please do not use any other chi-square tables to answer questions 1-3. Dearing et al. (2005)- I want you to answer the questions using Table 3 which is located on page 1399. Here is a screenshot Download screenshotof the table. Please do not use any other correlation tables to answer the questions. If you notice in the screenshot, I put a big red X through part of the table. This is because I only want you interpreting the bivariate correlations. Do not discuss/interpret part correlations in your responses. As a bonus, here is a helpful hint: Be sure you are reading the footnote carefully. * is significant at p<.05; ** is significant at p<.01; and *** is significant at p<0.001. Logically, if something is significant at the .01 level then it is also significant at the .05 level. Therefore, make sure you are interpreting all significant correlations, not just those that only significant at .05. Article 1: The selected article should contain Chi-square test(s) in a table or in text. If this article contains more than one Chi-square test, pick one and then answer the following questions. 1. What relationship is being tested in this Chi-square test? [2pts] Why is Chi-square test appropriate? [2pts] 2. How would you define this Chi-square test (e.g., 2x2, 2x3, etc.)? [1pt] 3. Interpret the result of this Chi-square test. In other words, did the authors find significant relationships between variables? Which relationships are significant, if any? What do these significant relationships tell you? [5pts] Article 2: 4. Answer the following questions about the correlations based on the table that the authors present bivariate correlations: a. Which is the strongest correlation in the entire table? Be sure to include the value of the correlation and describe the correlation (which variables are correlated). [2pts] 5. Which is the weakest correlation in the entire table? Be sure to include the value of the correlation and describe the correlation (which variables are correlated). [2pts] 6. Which variable(s) are consistently positively correlated with the criminological term you are interested in (e.g., crime, delinquency, sentencing length, etc.)? Interpret one positive relationship in sentence form (in the form of, as X increases, what happens to Y). [2pts] 7. Which variable(s) are consistently negatively correlated with the criminological term you are interested in (e.g., crime, delinquency, sentencing length, etc.)? Interpret one negative relationship in sentence form (in the form of, as X increases, what happens to Y). [2pts] 8. When researchers wish to report significant findings within a table, they often use asterisks. At the bottom of the table that presents bivariate correlations, you can see the note: * p < .05 for example. This note means that the authors performed a significant test on the correlations to determine if they had significant findings. They are telling us that all of the correlations with asterisks next to them were significant at the .05 alpha level. Using this information, answer the following questions about the correlations presented in the tables. a. Are there any correlations that are statistically significant? Be sure to include the value of the correlations and describe the correlation (which variables are correlated). [3pts] b. In this table, which variable(s) consistently demonstrate significant correlations with the criminological term you are interested in (e.g., crime, delinquency, sentencing length, etc.)? What do these significant correlations tell you? [4pts]See Answer
  • Q9:2- For the following synthetic data and quadratic model: import numpy as np import matplotlib.pyplot as plt np.random.seed (123) #Choose the "true" parameters for a quadratic model. a_true = 0.1 b_true = -1.0 c_true = 3.0 f_true = 0.5 # Generate some synthetic data from the quadratic model. N = 50/nx = np.sort (10* np.random.rand (N)) yerr 0.1 +0.5* np.random.rand (N) y = a_true * x**2 + b_true * x + c_true - y += np.abs (f_true * y) np.random.randn (N) y += yerr np.n random.randn (N) = * # Create a scatter plot with error bars plt.errorbar (x, y, yerr-yerr, fmt=".k", capsize=0) * # Generate the quadratic curve based on the "true" parameters x0 = np.linspace (0, 10, 500) y0 = a_true * x0**2 + b_true * x0 + c_true #Plot the quadratic curve plt.plot (x0, y0, "k", alpha=0.3, 1w=3) plt.xlim (0, 10) plt.xlabel ("x") plt.ylabel("y") # Show the plot plt.show()/ny 5 + 3 2 1 0- 0 2 #F₁ 8 10 Redo all the steps in the Generative Probabilistic Model section. Use these uniform priors for model parameters (0<a<3, -6<b<6, 2<c<5, and -2<f<2): (8 points)See Answer
  • Q10:1- In section "Generative Probabilistic Model", we used "w np. linalg. solve (ATA, np.dot (A.T, y / yerr**2)" to find the model parameters and fit a line to data. Use what we learned in chapter 2, least square fitting (page 7), to find the model parameters and fit a line to data. Provide your Python script and visualize the results like what we did in Chapter 2, page 8. (2 points) =See Answer
  • Q11: 1:46 Back Pulse Д Module Two Assignment Guidelines and Rubric Listen MAT 240 Module Two Assignment Guidelines and Rubric 91 LB > Scenario Smart businesses in all industries use data to provide an intuitive analysis of how they can get a competitive advantage. The real estate industry heavily uses linear regression to estimate home prices, as cost of housing is currently the largest expense for most families. Additionally, in order to help new homeowners and home sellers with important decisions, real estate professionals need to go beyond showing property inventory. They need to be well versed in the relationship between price, square footage, build year, location, and so many other factors that can help predict the business environment and provide the best advice to their clients. Prompt You have been recently hired as a junior analyst by D.M. 1:46 91 Back Prompt Pulse You have been recently hired as a junior analyst by D.M. Pan Real Estate Company. The sales team has tasked you with preparing a report that examines the relationship between the selling price of properties and their size in square feet. You have been provided with a Real Estate Data Spreadsheet spreadsheet that includes properties sold nationwide in recent years. The team has asked you to select a region, complete an initial analysis, and provide the report to the team. Note: In the report you prepare for the sales team, the response variable (y) should be the listing price and the predictor variable (x) should be the square feet. Specifically you must address the following rubric criteria, using the Module Two Assignment Template: • Generate a Representative Sample of the Data • Select a region and generate a simple random • sample of 30 from the data. Report the mean, median, and standard deviation of the listing price and the square foot variables. Analyze Your Sample 。 Discuss how the regional sample created is or is о not reflective of the national market. ■ Compare and contrast your sample with the population using the National Summary Statistics and Graphs Real Estate Data PDF document. 。 Explain how you have made sure that the sample is random. ■ Explain your methods to get a truly random 1:46 Back Pulse 91 Analyze Your Sample 。 Discuss how the regional sample created is or is not reflective of the national market. ■ Compare and contrast your sample with the population using the National Summary Statistics and Graphs Real Estate Data PDF document. 。 Explain how you have made sure that the sample is random. ■ Explain your methods to get a truly random sample. • Generate Scatterplot • Create a scatterplot of the x and y variables noted above. Include a trend line and the regression equation. Label the axes. • Observe patterns 。 Answer the following questions based on the scatterplot: ■ Define x and y. Which variable is useful for making predictions? ■ Is there an association between x and y? Describe the association you see in the scatter plot. ■ What do you see as the shape (linear or nonlinear)? ■ If you had a 1,800 square foot house, based on the regression equation in the graph, what price would you choose to list at? ■ Do you see any potential outliers in the scatterplot? ■ Why do you think the outliers appeared in the scatterplot you generated? ■ What do they represent? 1:46 91 Back Pulse regression equation. Label the axes. • Observe patterns • Answer the following questions based on the scatterplot: ■ Define x and y. Which variable is useful for making predictions? ■ Is there an association between x and y? Describe the association you see in the scatter plot. ■ What do you see as the shape (linear or nonlinear)? ■ If you had a 1,800 square foot house, based on the regression equation in the graph, what price would you choose to list at? Do you see any potential outliers in the scatterplot? Why do you think the outliers appeared in the scatterplot you generated? ■ What do they represent? You can use the following tutorial that is specifically about this assignment. Make sure to check the assignment prompt for specific numbers used for national statistics and/or square footage. The video may use different national statistics or solve for different square footage values. • MAT-240 Module 2 Assignment You can also use the following tutorials for support as you develop the report: Load the Analysis ToolPak in Excel 1:47 Back Pulse 91 different national statistics or solve for different square footage values. MAT-240 Module 2 Assignment You can also use the following tutorials for support as you develop the report: • Load the Analysis ToolPak in Excel . Downloading Office 365 Programs PDF • Random Sampling in Excel PDF • Scatterplots in Excel PDF • Descriptive Statistics in Excel PDF What to Submit Submit your completed Module Two Assignment Template as a Word document that includes your response, supporting charts, and Excel file. Module Two Assignment Rubric Generate a Representative Sample of the Data Exemplary N/A Proficient Includes a random sample of 30 from a region and descriptive statistics for the sample (100%)See Answer
  • Q12: EPOM405 Homework Assignment M9 1. Use the following methods to test Ho: μ₁ = μ2 versus H₁ μ₁ ± μ2 for the data in worksheet 9.1 of the Excel data file: a. Tukey's Quick Test b. A boxplot slippage test c. The two-sample t test d. The Mann-Whitney test (use Stat> Nonparametrics> Mann-Whitney) 2. Use the following methods to test H0 µ₁ = μ2 versus H₁ : µ₁ ± µ2 for the data in worksheet 9.2 of the Excel data file: a. Tukey's Quick Test b. A boxplot slippage test c. The two-sample t test d. The Mann-Whitney test 3. Use the data from worksheet 9.2 to test for a difference between the two population standard deviations: a. By manually calculating the F statistic and comparing it to the F0.95 critical value. b. By calculating the p value for the F statistic (Cal> Probability Distributions or Graph> Probability Distribution Plot) and comparing it to a = = 0.05. c. Using the F statistic in MINITAB via Stat> Basic Statistics> 2 Variances> Options> Use Test and Confidence Intervals Based on Normal Distributions. d. Using Bonnet's and Levene's tests in MINITAB by turning off the normality assumption option in part c. e. Compare the results of the F, Bonnet's, and Levene's methods. Do they agree? What if they don't? 4. Determine the 95% confidence interval for the true process cp if a sample of size n = S = 0.0042 and the spec is USL/LSL = 1.500 ± 0.020. 80 gives 5. An experiment was performed to study the defective rate of a process. The hypotheses to be tested were Ho : p 0.01 vs. H₁ : p > 0.01. A sample of size n = 100 parts was drawn, inspected, and = found to have D = 3 defectives. (Hint: Use Stat> Basic Statistics> 1 Proportion.) a. Is there sufficient evidence to reject Ho? b. What is the one-sided upper 95% confidence limit for p? 6. In an experiment to compare the incidence of heart disease in populations eating the Mediterranean diet vs. a western diet 14 of 219 people from the Mediterranean diet group had heart attacks and 44 of 204 people from the western diet group had heart attacks. a. Use Fisher's Exact Test and the large sample (i.e. normal approximation to the binomial distribution) test to determine if there is a difference in the heart attack rates. (Hint: Use Stat> Basic Statistics> 2 Proportions.) b. Use the x² test for association to determine if there is a difference in the heart attack rates. (Hint: Use Stat> Basic Statistics> Tables> Cross Tabulation and Chi-square or > Chi-square Test for Association.) c. Do the three methods agree? When will the Fisher's method be preferred and the other two methods avoided? MM&B Inc. 7. Precision valves for controlling the flow of fluids and liquids can be assembled in a clean room or on a less-clean production line. The cleanliness of the assembly process can affect the leak rate. Twenty valve assemblies were made on the production line and another twenty were made in the clean room. The same component lots were used for all forty assemblies. a. There were three leakers from the production line versus zero leakers from the clean room. Is there sufficient evidence to conclude that the leak rate on the production line is worse than the leak rate in the clean room? b. What's the smallest number of defectives that would have to be found on the production line with no defectives found in the clean room to conclude that the production line leaker rate is worse that in the clean room? 8. Cheap plastic children's dice are made by pressing divots in the die faces causing the dice to be unbalanced, e.g. the side with a 6 will be light and the side with a 1 will be heavy. To test this claim, dice were rolled 180 times and the observed frequencies of 1-6 were found to be 30, 32, 22, 34, 39, 23. (The data are in worksheet 9.8.) Test the observed frequencies to see if they are consistent with the null hypothesis Ho: The dice are balanced versus HA: The dice are unbalanced by: a. Manual calculate the x² statistic and compare it to the appropriate critical value. b. Confirm your answer to part a using MINITAB Stat> Tables> Chi-square Goodness of Fit Test (One Variable). MM&B Inc. EPOM405 Homework Assignment M9 Answers to Selected Questions 1. a) (T = 6) < (T0.05 = 7) so cannot reject Ho b) boxes are overlapped, cannot reject Ho c) t = −2.16, p = 0.059 so cannot reject Ho at a 0.05 d) (p = 0.071) > (α = 0.05) so cannot reject Ho 2. a) T = 7 so reject Ho b) sample sizes are large and one median is slipped from the other sample's box so reject Ho c) t = -2.78, p = 0.008 so reject Ho at a = 0.05 d) p = 0.0023 so reject Ho at a = 0.05 3. a. (F = 2.72) > (F0.95 = 1.76) so reject Ho and conclude that the variances are not equal b. One-tailed test, (p = 0.0019) < (α = 0.05) so reject Ho and conclude that the variances are not equal c. MINITAB's F two-tailed test d. PBonnet = = 0.37 is the reciprocal of F = 2.72 and its p = 0.004 value is for its default = 0.023, both two-tailed tests 0.008 and p Levine e. All of the methods agree 4. P(1.38 < Cp < 1.80) = 0.95 5. a. p 0.079 b. LCL = 0.0082 6. a. pFisher = 0.000, z = -4.57, p = 0 7. b. x² 20.6 with df = 1, p = 0 a. p = 0.115 b. 5 8. x² = 7.13, p = 0.211 MM&B Inc.See Answer
  • Q13:(b) Calculate a 95% two-sided confidence interval on 01/02. Round your answer to three decimal places (e.g. 98.765). VI 56 VI i 02 Statistical Tables and Charts eTextbook and Media Save for Later Attempts: 0 of 3 used SUPPORT for bothSee Answer
  • Q14:Problem 2. Medical test (#Probability, #MathTools) Omer wants to be really certain about a diagnosis so he takes a series of identical medical tests. He hopes multiple tests will reduce his uncertainty. Events: D: Omer has the disease being tested for. T; : Omer tests positive on the jth test for j = 1, 2, ..., n. Let p = P(D) be the prior probability that he has the disease. Question 4 of 15 2.1 Assume for this part the test results are conditionally independent given Omer's disease status. Let a = P(T;|D) and b₁ = P(T;|D), where ao and bo don't depend on j. Find the posterior probability that Omer has the disease, given that he tests positive on all of the tests. Hint: Since a does not depend on j (the index of a test result), the algebra simplifies a lot since P(T;|D)P(T;|D) = a for any test indices i and j. The same holds for bn. Question 5 of 15 2.2 Suppose some people have a gene that makes them always test positive on this type of medical test. Let G be the event that Omer has the gene. Assume that P(G) and that D and G are independent - that is, the gene does not make you more or less susceptible to the disease. If Omer has the gene, he'll test positive on all tests. If Omer does not have the gene, then the test results are conditionally independent given his disease status. Let a₁ = P(T;|D,G) and b₁ = P(T;|Dº‚Gº), where a₁ and b₁ don't depend on j. Now, suppose that Omer tests positive on all tests and find the posterior probability that Omer has the disease. Question 6 of 15 2.3 Using the same setup as in part (2.), find the posterior probability that Omer has the gene given that he tests positive on all n of the tests.See Answer
  • Q15:Problem 1. Coin Spins (#Probability, #MathTools) One day you overhear how two fellow ICS students discuss a show in which a performer spins a coin many times, resulting in Tails 30 times in a row. You hear them discussing: Yueh Han: "That outcome would be extremely unlikely with fair coins. They must be using trick coins (maybe double-tailed coins), or the experiment must have been rigged somehow (maybe with magnets)." Akma: "It's true that the string TTT of length 30 is very unlikely; the chance is with fair coins. But any other specific string of H and T with length 30 has exactly the same probability! The reason the outcome seems extremely unlikely is that the number of possible outcomes grows exponentially as the number of spins grows, so any outcome would seem extremely unlikely. You could just as well have made the same argument even without looking at the results of their experiment, which means you really don't have evidence against the coins being fair Help Yueh Han and Akma resolve their debate. Question 1 of 15 1.2 Suppose there are only two models: either the coins are all fair (and spun fairly), or double-tailed coins are being used in which case the probability of Tails is 1. Let p be the prior probability that the coins are fair. Find the posterior probability that the coins are fair, given that they landed Tails in 30 out of 30 trials. 1.3 For which values of the prior, p, is the posterior probability that the coins are fair greater than 0.5? (What does the prior need to be to make the fair-coin model more probable according to the posterior?)See Answer
  • Q16:3 Lakers This question uses data on basketball games involving the LA Lakers in the 2008-2009 season. Once you've loaded in the tidyverse package, you should be able to access the data using data("lakers"). The outcome of interest is whether or not a particular shot i made the basket: 0 if shot i missed. We are interested in studying the association between the shot being made and where the shot was taken on the court. The variables of interest for these questions are • result, which you can use to create z; above; is the horizontal coordinate of where the shot was taken; • y, which is the vertical coordinate of where the shot was taken. If you do a search ?lakers, this will tell you a bit more about the dataset. If you go to the original source (www.basketballgeek.com/data/), this tells you a bit more about how to interpret the (x,y) coordinates. Note that only shots have (x,y) coordinates, so for this question you can filter out all other events (rebounds, free throws, etc). a) Do an exploratory data analysis illustrating the relationship between making a shot, and the location on the court where the play was made. Think about different ways of effectively illustrating the relationships given the binary outcome. As usual, a good EDA includes well- thought-out descriptions and analysis of any graphs and tables provided, well-labelled axes, titles etc. Assume z; ~ Bern(p;), where p; refers to the probability of making a shot Consider two candidate models. . Model 1: • Model 2: logit (pi) = Bo + B₁ (Ti — Io) + B₂ · (Yi - Yo) + ß3· (Ti — xo) (Yi - Yo) logit (p;) = 30 +3₁ (|T₁ – 1o|) + B₂ · (Yi - Yo) + B3 · (|Ti — xo) (Yi - Yo) where z; is the x-coordinate and y, is the y-coordinate of the shot. The values zo and yo refer to the coordinates of the basket./nwhere z; is the x-coordinate and y; is the y-coordinate of the shot. The values zo and yo refer to the coordinates of the basket. b) Fit both of these models using Stan. Put N(0, 1) priors on all the 3s. You should generate pointwise log likelihood estimates (to be used in later questions), and also samples from the posterior predictive distribution (unless you'd prefer to do it in R later on). For both models, interpret each coefficient. 4 c) Let t(z) = 1 1 (zi= 1, yi > 10)/₁1 (>10) i.e. the proportion of shots made at a y-distance greater than 10. Calculate t(zrep) for each replicated dataset for each model, plot the resulting histogram for each model and compare to the observed value of t(z). Calculate P (t (zep) <t(z)) for each model. Interpret your findings. d) Use the loo package to get estimates of the expected log pointwise predictive density for each point, ELPD₁. Based on Σ; ELPDi, which model is preferred? e) Create a scatter plot of the ELPD;'s for Model 2 versus the ELPD's for Model 1. Create another scatter plot of the difference in ELPD;'s between the models versus the (centered) y-coordinate. In both cases, color the dots based on the value of zi. Interpret both plots. f) Given the outcome in this case is discrete, we can directly interpret the ELPD;s. In particular, what is exp(ELPD;)? g) For each model recode the ELPD;'s to get 2 = E(Z; z-i). Create a binned residual plot, looking at the average residual zi - 2 by (centered) y-coordinate. Split the data such that there are 40 bins. On your plots, the average residual should be shown with a dot for each bin. In addition, add in a line to represent +/- 2 standard errors for each bin. Interpret the plots for both models.See Answer
  • Q17:Part D: *****MUST DO QUESTION***** 1 The net incomes of a sample of 20 container shipping companies were organized into the following table: Net Income ($ millions) Number of Companies 2 up to 6 1 6 up to 10 10 up to 14 14 up to 18 18 up to 22 a. What is the table called? 4 10 d. Calculate the Median, Q3, P45 and Range 3 2 b. Based on the distribution, what is the estimate of the arithmetic mean net income? c. Based on the distribution, what is the estimate of the standard deviation?See Answer
  • Q18:/n collective_efficacy punitiveness 1.333255887 3.841360807 2.81778264 2.75338912 1.843444943 1.848545909 1.341444731 2.065723896 3.657845497 1.629728198 2.982000113 1.498682857 2.843961239 2.396447897 2.790726662 1.717732787 1.619872093 2.947984219 2.304697752 3.646073103 2.153009176 1.569896579 2.613511324 2.712970495 3.579788446 2.718465805 1.079171896 2.96315527 1.193508625 3.976508379 1.950628281 3.310344696 1.573178649 3.558977127 1.105786324 3.937344551 1.363484025 1.682451963 1.709320664 1.614470005 3.94071579 2.641096354 3.524244547 2.517956257 2.648793936 3.096044302 3.927428722 2.182617188 1.341046691 1.980771899 3.824162245 3.667945862 1.208176136 1.150792599 1.668160439 3.238552809 3.479876041 3.567532539 1.589430451 3.750486612 1.567864299 1.61555028 2.373256922 3.717535734 2.105895281 1.612153888 2.82738018 3.252717733 3.116099834 1.422777653 3.48306489 1.594265223 2.248568296 3.86125946 1.87931025 3.764187574 3.234582663 3.800332308 2.73434329 2.721946239 2.205566406 3.799867868 1.775756598 3.698415279 1.798564672 2.834188223 PUNITIVENESS 3.088351965 1.994250059 2.308962345 1.186883926 2.948550463 1.649141788 1.240551353 3.960311413 1.56605804 2.75984621 2.735659838 2.764545679 2.848798037 1.701640129/n person 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 sex 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 race age 1 1 2 2 2 2 2 1 1 2 2 2 2 2 2 1 2 1 2 2 2 2 2 2 1 1 1 2 2 2 crime 52 43 22 47 63 31 76 70 37 49 54 68 23 47 61 33 45 34 54 29 56 29 30 21 51 55 78 72 46 42 recidivism custody 1 Minimum 0 Close 0 Close O Close 0 Minimum 1 Medium 1 Medium 0 Close 1 Medium 1 Medium 0 Minimum 0 Minimum 1 Medium 0 Close 0 Medium 0 Close 0 Minimum 0 Minimum 1 Minimum 0 Medium 0 Medium 1 Medium 1 Minimum 0 Medium 0 Medium 1 Medium 1 Medium 0 Medium 0 Medium 0 Medium 2 3 1 1 2 3 2 1 2 2 2 2 1 3 2 2 3 1 2 2 2 3 1 1 1 2 2 3 2 2 RECID IVISM/n Gray (2010)- I want you to answer the questions using Table III which is located on page 547. Here is a screenshot Download screenshot of the table. Please do not use any other chi-square tables to answer questions 1-3. Problem behavior Low Marijuana use (lifetime) Recent excessive drinking Medium Arrest Drugs other than marijuana (lifetime) High Recent marijuana use Misdemeanor conviction Theft in previous 12 months Burglary in previous 12 months Sold drugs in previous 12 months Recent other drug use Felony conviction Notes: *p<0.05; **p < 0.01 Policing (n = 171) % 51.5 33.9 11.7 14.0 15.2 9.4 5.3 4.1 4.1 1.2 0.6 Others (n = 703) % 53.3 26 15.1 22.5 15.9 13.7 5.4 2.6 6.5 4.6 1.7 Chi-square 0.20 4.30* 1.30 5.96** 0.05 2.30 0.01 1.16 1.45 4.20* 1.19 Problem behaviors of students 547 Table III. Problem behaviors Dearing et al. (2005)- I want you to answer the questions using Table 3 which is located on page 1399. Here is a screenshot. Please do not use any other correlation tables to answer the questions. If you notice in the screenshot, I put a big red X through part of the table. This is because I only want you to interpret the bivariate correlations. Do not discuss/interpret part correlations in your responses. Table 3 Study 3 relations of shame-proneness and guilt-proneness to alcohol and drug use and problems Problems/use Mean SD Bivariate correlations Rart correlations Shame Guilt Shame Alcohol PAI alcohol problems TCU-CRTF-frequency of use TCU-CRTF-dependence Drug PAI drug problems TCU-CRTF-frequency of use Cocaine Marijuana Polydrug TCU-CRTF-Dependence Cocaine Marijuana 11.40 3.41 .80 14.63 1.78 2.36 1.95 .91 .63 10.02 2.45 1.05 9.91 2.57 2.78 1.99 1.43 1.05 .12* -.02 .17** .21*** .17** -.08 .11* .21*** .14** -.04 -.03 -.06 -.04 .04 -.19*** -.13* .04 -.14** .13 -.01 .18** .22*** .16** -.04 4** .20*** .18*** Guilt -.07 -.03 -.10t -.09 -.00 -.18*** .16** -.01 -.18*** N=307-332. PAI=Personality Assessment Inventory; TCU-CRTF=Texas Christian University Correctional: Residential Treatment Form, Initial Assessment. *p<.05; **p<.01; ***p<.001; †=.06. Be sure you are reading the footnote carefully. * is significant at p<.05; ** is significant at p<.01; and *** is significant at p<0.001. Logically, if something is significant at the .01 level then it is also significant at the .05 level. Therefore, make sure you are interpreting all significant correlations, not just those that only significant at .05. 1. Use the Recidivism.xls data, conduct a Chi-Square independence test (alpha= .05) in Excel to determine whether there is a significant relationship between RACE (1=Black, 2=White) and RECIDIVSM (1=the ex-offender recidivated, 0=the ex-offender did not recidivate). Copy and paste the Excel output here. In your summary, please discuss the null hypothesis, alternative hypothesis, the result of this hypothesis test, and interpret the result. 2. Use the Punitiveness.xls data, correlate collective_efficacy (respondents' perceived collective efficacy in their neighborhood) and punitiveness (respondents' support for punitive criminal justice policies) in Excel. Copy and paste the Excel output here. Summarize your results and make sure that you briefly explain what this correlation means to you in terms of the impact of collective efficacy on respondents' support for punitive criminal justice policies (direction of the correlation, correlation coefficient, and test of significance result).See Answer
  • Q19:A Pew Research Center survey (Pew Research website) examined the use of social media platforms in the United States. The survey found that there is a 0.63 probability that a randomly selected American will use Instagram and a 0.34 probability that a randomly selected American will use LinkedIn. In addition, there is a 0.26 probability that a randomly selected American will use both Instagram and LinkedIn. a) What is the probability that a randomly selected person will use Instagram or LinkedIn or both (round your answer upto 2 decimals)? (HINT: Use Addition Law) b) What is the probability that a randomly selected person will not use either social media platform (round your answer upto 2 decimals)? (HINT: Use the relationship between probability of an event and probability of complement of an event) Click Save and Submit to save and submit. Click Save All Answers to save all answers.See Answer
  • Q20:Please complete the exercise 7.9 "Excel MLR Practice: Disney Movie Revenue" found https://app.myeducator.com/reader/web/1382fc/chp07/kz4jl/ for Chapter 7 assignment, but in R instead of Excel. To do this, download the file disney_movies_total_gross.csv from the book under Resource Files and load it into RStudio. As you work on each question from the book, post your answer in the book. Here are the steps in more detail: 1. Download the file disney_movies_total_gross.csv from the book under Resource Files. Open RStudio and load the data file. 2. Once the data file is loaded, you can start working on the questions in the exercise. 3. To answer each question, write R codes to perform the necessary tasks and then post your answer in the book. As you work on the assignment, please take a look at the lecture (5 minutes long, link provided below) on how to interpret the coefficients for categorical variables with more than two levels in regression models. Please log into MyJSU before clicking on the link. https://www.linkedin.com/learning/introduction-to-stata-15/categorical-explanatory-variables-in- ols?u=36441276 The lectures I posted on Canvas explain how to interpret a model result mostly using continuous variables, except for the Gender variable, which has two levels (0/1). The lecture directed by the link above explains the output of a model with a categorical variable with more than four levels. This lecture will help you to interpret the coefficients of your model in the assignment./nDeliverables: 1. Please create a separate Word file that will include an interpretation of your model results for Question #7. Specifically, please interpret the results for the coefficients, Std. Error, t-value, p- value for each of the variables (days_since_release, mpaa_rating) produced in R. My lectures posted on Canvas can help you with the answers to this question. 2. For question 17, upload your R file to Canvas, with your code separated by question number and with the question number included before the code for that question. 1/n7) Fill in all missing values of mpaa_rating with the value "Empty" Create dummy code features to convert mpaa_rating to sets of numeric 0/1 features for all values except "Empty" Run another MLR a few rows below the previous one using both days_since_release and all of the dummy codes you just created for mpaa_rating What did the inclusion of these dummy coded mpaa_rating features do to the model fit?/n17)Upload the Excel file containing the data and all of your regression models 7,17questions and read the instructions for both questions in assignment document. 7,17 give above are questions for above mentioned exercise 7.9 Excel MLR Practice: Disney Movie Revenue Objective Create an MLR model to explain and predict the gross revenue of Disney movies from 1937 to 2016. Data Source Use the .csv file provided below, which includes 573 records with the following features: Labels total_gross: the actual gross revenue of the movie inflation_adjusted_gross: the gross revenue converted to account for inflation Features movie_title: the title of the film release_date: the first date it appeared in theaters genre: the type of file mpaa_rating: G, PG, PG-13, R, Not Rated, or null/empty Tasks Perform the steps included in each of the questions below and answer the associated questions. DeliverablesSee Answer

TutorBin Testimonials

I found TutorBin Advanced Statistics homework help when I was struggling with complex concepts. Experts provided step-wise explanations and examples to help me understand concepts clearly.

Rick Jordon

5

TutorBin experts resolve your doubts without making you wait for long. Their experts are responsive & available 24/7 whenever you need Advanced Statistics subject guidance.

Andrea Jacobs

5

I trust TutorBin for assisting me in completing Advanced Statistics assignments with quality and 100% accuracy. Experts are polite, listen to my problems, and have extensive experience in their domain.

Lilian King

5

I got my Advanced Statistics homework done on time. My assignment is proofread and edited by professionals. Got zero plagiarism as experts developed my assignment from scratch. Feel relieved and super excited.

Joey Dip

5

TutorBin helping students around the globe

TutorBin believes that distance should never be a barrier to learning. Over 500000+ orders and 100000+ happy customers explain TutorBin has become the name that keeps learning fun in the UK, USA, Canada, Australia, Singapore, and UAE.