So, when you first look at some data, it’s helpful to get a feel of it. One way to do this is to do a plot or two. I’ve found myself continuously doing the same series of plots for different datasets, so in the end I wrote this short code to put all the plots together as a time saving device. Not pretty, but gets the job done.
The output looks like this:

So on the top a histogram with a normal distribution plot. On the right a QQ normal plot, with an Anderson Darling p value. Then in the middle on the left is the same data put into different numbers of bins, to see how this affects the look of the data. And on the right, we pretend that each value is the next one in a time series with equal time intervals between readings, and plot these. Below this is the ACF and PACF plots.
Hope someone else finds this useful. If there’s easier ways to do this, let me know. To use the code – put your data into a text file as a series of numbers called data.txt in the working directory, and run this code:
01 |
## univariate data summary |
03 |
data <- as.numeric(scan ("data.txt")) |
04 |
# first job is to save the graphics parameters currently used |
05 |
def.par <- par(no.readonly = TRUE) |
06 |
par("plt" = c(.2,.95,.2,.8)) |
07 |
layout( matrix(c(1,1,2,2,1,1,2,2,4,5,8,8,6,7,9,10,3,3,9,10), 5, 4, byrow = TRUE)) |
09 |
#histogram on the top left |
10 |
h <- hist(data, breaks = "Sturges", plot = FALSE) |
11 |
xfit<-seq(min(data),max(data),length=100) |
12 |
yfit<-yfit<-dnorm(xfit,mean=mean(data),sd=sd(data)) |
13 |
yfit <- yfit*diff(h$mids[1:2])*length(data) |
14 |
plot (h, axes = TRUE, main = "Sturges") |
15 |
lines(xfit, yfit, col="blue", lwd=2) |
16 |
leg1 <- paste("mean = ", round(mean(data), digits = 4)) |
17 |
leg2 <- paste("sd = ", round(sd(data),digits = 4)) |
18 |
legend(x = "topright", c(leg1,leg2), bty = "n") |
21 |
qqnorm(data, bty = "n", pch = 20) |
24 |
leg <- paste("Anderson-Darling p = ", round(as.numeric(p[2]), digits = 4)) |
25 |
legend(x = "topleft", leg, bty = "n") |
27 |
## boxplot (bottom left) |
28 |
boxplot(data, horizontal = TRUE) |
29 |
leg1 <- paste("median = ", round(median(data), digits = 4)) |
30 |
lq <- quantile(data, 0.25) |
31 |
leg2 <- paste("25th quantile = ", round(lq,digits = 4)) |
32 |
uq <- quantile(data, 0.75) |
33 |
leg3 <- paste("75th quantile = ", round(uq,digits = 4)) |
34 |
legend(x = "top", leg1, bty = "n") |
35 |
legend(x = "bottom", paste(leg2, leg3, sep = "; "), bty = "n") |
37 |
## the various histograms with different bins |
38 |
h2 <- hist(data, breaks = (0:12 * (max(data) - min (data))/12)+min(data), plot = FALSE) |
39 |
plot (h2, axes = TRUE, main = "12 bins") |
41 |
h3 <- hist(data, breaks = (0:10 * (max(data) - min (data))/10)+min(data), plot = FALSE) |
42 |
plot (h3, axes = TRUE, main = "10 bins") |
44 |
h4 <- hist(data, breaks = (0:8 * (max(data) - min (data))/8)+min(data), plot = FALSE) |
45 |
plot (h4, axes = TRUE, main = "8 bins") |
47 |
h5 <- hist(data, breaks = (0:6 * (max(data) - min (data))/6)+min(data), plot = FALSE) |
48 |
plot (h5, axes = TRUE,main = "6 bins") |
50 |
## the time series, ACF and PACF |
51 |
plot (data, main = "Time series", pch = 20) |
52 |
acf(data, lag.max = 20) |
53 |
pacf(data, lag.max = 20) |
55 |
## reset the graphics display to default |