Introduction to the Assignment
Hello everyone and welcome to your final assignment for the course! This assignment asks you to apply several of the methods and concepts that you've learned in the course, to the analysis of some data collected in a scientifc study. You'll have to look through various tables and figures about these data, and answer six different statistical questions about them, using the statistical ideas you have learned throughout this MOOC.
Some of the questions on this assignment require you to do some calculation. All of the calculations can be carried out using only a hand calculator. The tests and confidence intervals on this assignment are such that you do not need to use the computer or a table to get values from a normal or
Note that this assignment has a HARD DEADLINE this coming Sunday (May 26) at 11:59 PM EDT, with no flexibility about that (sorry) because we need to start the peer assessment at that time. If you do not enter work before the submission deadline, you will not receive an evaluation for this assignment, nor will you be allowed to evaluate your classmates' work. So, please get started early, and don't forget to SUBMIT your work before the deadline.
You can create and save multiple drafts any time during the submission process; your draft enters the system to be graded only when you click the "Submit for grading" button. (This means that if you haven't clicked "Submit for grading" by the deadline, your saved draft will not be evaluated, so please be sure to click it!) Even after submitting, you can return to your assignment, make changes and "Re-submit for grading" anytime up until the deadline. You will need to check the box "In accordance with the Honor Code..." to activate the Submit button.
Next week, we will continue with the peer assessment phase of this assignment. You will be asked to review five other students' assignments, and five other students will review your assignment too. Some general information about peer assessment is available here, and we will provide you with more details next week. Remember that you must participate fully in BOTH the assignment AND the peer assessment in order to receive credit for your work.
Background for the Assignment
In the week 8 lectures, we looked at some data from the Canadian Trial of Carbohydrates in Diabetes (CCD). For the CCD study, the primary objective was to investigate how HbA1c, a measure of diabetic controlled, differed between subjects on a low glycemic index diet and subjects on a high glycemic index diet, where the subjects all had type 2 diabetes which was being treated with diet alone. Like most large studies, much more data were collected than what was needed to study the primary objective, and many secondary analyses were carried out. In this assignment, we will consider part of a secondary analysis of some of the data that were collected on the subjects at baseline, that is before the subjects started their treatments.
The key pdf document for this assignment is posted here. That document gives more information about the data being considered, including several tables and figures (plots). For this assignment, please use ONLY the information provided in that document. (We won't be providing the data for now, partially because we don't want to disadvantage students who are not learning R.)
The Actual Assignment
Your assignment is to use the information in the pdf document of tables and figures mentioned above to answer the following six questions, using the material that we have covered in this MOOC in Weeks 1 through 7. Please try to be concise and to the point in your answers.We're interested in the prevalence of the A variant of the TNF-alpha gene at position 308. It is thought that this variant is less common, that is, it occurs in less than half of the population. Carry out an hypothesis test to test this. In your answer, state:
- Whether you're carrying out a one-sided or two-sided test.
- The value of the test statistic that you calculate.
- Based on the test statistic, whether the P-value is large or small.
- Your conclusion.
From our data, we estimate that 28.2% of the population has the A variant at position 238, and 30.0% has the A variant at position 308, where the "population" here is the population that is represented by our subjects. We're interested in whether the prevalence of the A variant is the same for position 238 and position 308 in this population. That is, is the proportion of people with the A variant at position 238 the same as the proportion of people with the A variant at position 308? Using only the information you are given in the pdf document of plots and summary statistics, this test cannot be carried out. Why not?
Construct a 95% confidence interval for the difference in the proportion of males with the A variant at position 308, and the proportion of females with the A variant at position 308. State the confidence interval that you calculate. Based on your confidence interval, is there evidence that the proportion of the A variant at 308 differs between males and females? You may use the fact that the critical value for a 95% confidence interval from a standard normal distribution is 1.96.
Carry out a statistical test to determine if the mean of HDL is the same for males and females for the population represented by this study. You should assume that the variance is the same for males and females. In your answer, state:
- Whether you're carrying out a one-sided or two-sided test.
- The value of the test statistic that you calculate.
- Based on the test statistic, whether the P-value is large or small.
- Your conclusion.
It is believed that higher consumption of polyunsaturated fats (PUFA) increases HDL cholesterol levels, but this relationship may be affected by whether or not a subject has the A genetic variant at position 308. Does this seem to be the case based on the plots and summary statistics you are given? Investigate this question by examining whether or not there is a relationship between HDL and PUFA, and whether or not that relationship differs between subjects who have and do not have the A variant at position 308. Support your answer by mentioning relevant plots or summary statistics. Indicate whether or not you have any concerns about the appropriateness of the analysis of the HDL - PUFA relationship; that is, might a different type of analysis, or analysis on modified data be more appropriate?
In addition to polyunsaturated fats, many different components of diet could be investigated (for example, total calories consumed, and percent of dietary intake from protein, carbohydrates, alcohol, and other types of fat), and many different risk factors of heart disease could be investigated (for example, LDL cholesterol, triglycerides, total cholesterol, and C-reactive protein). So this assignment reflects only one small part of the analysis that was carried out. How does knowing this affect the implications of the conclusions you made?

