Biostatistics-CDFD-2026
A Biostatistics course offered to PhD students at BRIC-Centre for
DNA Fingerprinting and Diagnostics - CDFD during August-December
2026.
Lecture notes
Lecture 01 introduces R:
- R Basics — basic syntax, variables, vectors, data
frames, descriptive statistics, plots, statistical tests, correlation,
and linear regression.
- Reading data files — reading CSV/TSV files,
inspecting data, selecting rows and columns, modifying data, and saving
results.
- ggplot2 — introduction to plotting with
ggplot2, including scatter plots and fitted regression
lines.
- Bioconductor — a first example using
Biostrings to work with DNA sequences.
Lecture 02 introduces discrete probability distributions and
hypothesis testing using biological examples:
- Discrete random variables — introduction to
discrete data, probability distributions, and biological examples
involving counts.
- Binomial distribution — modeling repeated
independent trials with two possible outcomes, illustrated using coin
tosses and implemented in R with
dbinom() and
pbinom().
- Poisson distribution — modeling counts of rare
events and understanding the Poisson approximation to the binomial
distribution when
is large and
is small.
- HIV mutation example — modeling the number of
mutations in an HIV genome using the binomial distribution and comparing
it with the Poisson approximation.
- Cumulative probability — calculating probabilities
such as
by summing individual probabilities and using cumulative distribution
functions.
- Hypothesis testing — introducing the null and
alternative hypotheses, p-values, significance levels, rejection
regions, and left-tailed tests using the HIV mutation example.
- Confidence and significance levels — connecting
with a corresponding 95% probability level and interpreting significance
thresholds in discrete probability distributions.
Lecture 03 develops exact binomial hypothesis testing using
biological examples:
- When to use the binomial distribution — recognizing
situations with a fixed number of independent trials, two possible
outcomes per trial, and a constant probability of success.
- Null model and observed proportion — distinguishing
between the null probability
,
the unknown underlying probability
,
the observed proportion
,
and the expected count
.
- Hypothesis testing and compatibility regions —
interpreting
as defining an approximately 95% central region of outcomes expected
under
and rejection regions in the tails.
- p-values and the CDF — calculating left- and
right-tailed p-values using cumulative probabilities with
pbinom() and understanding the connection between the CDF
and hypothesis testing.
- Exact binomial test in R — using
binom.test() for left-tailed, right-tailed, and two-sided
tests and extracting the p-value using result$p.value.
- Mendel’s pea experiment — testing whether the
observed proportion of round seeds differs from the Mendelian
expectation of
.
- Allele-specific expression — testing whether
allele-specific sequencing reads deviate from the 50:50 expectation
using a two-sided exact binomial test.
- EXACT precision-oncology trial — performing a
right-tailed exact binomial test with
and connecting the p-value with the critical rejection region.
Additional content will be added as the course
progresses.
Repository
The repository https://github.com/raghurama123/Biostatistics-CDFD-2026
contains R scripts, data, and notes.
The repository is organized lecture-wise. Each lecture folder
contains the corresponding R scripts, data files, and lecture notes.
Biostatistics-CDFD-2026/
│
├── README.md
│
└── Lec01/
| ├── Lec01.md
| ├── Lec01.html
| ├── Rbasics.R
| ├── ...
└── Lec02/
| ├── ...
Interactive plots
How to do various
things in R programming?
The HowToR folder contains short
manuals and examples explaining how to perform various tasks in
R.
General reference
The main general reference for the course is the online book:
Modern Statistics for
Modern Biology by Susan Holmes and Wolfgang Huber
The book provides a modern introduction to statistical thinking and
data analysis for biological applications, with extensive use of R.
Raghunathan Ramakrishnan
Tata Institute of Fundamental Research Hyderabad, India
Email: ramakrishnan@tifr.res.in