How do I get started on R and R Studio?

This post is intended as a non-comprehensive simple guide to help existing psychologists who are trained in SPSS to replicate some of those skills in R via R Studio.

Using two worked examples, it introduces you to how R Studio operates and to the process of writing and running analytical code. The beauty of R is that you can tailor your analyses in a million different ways. Don’t presume these examples even scratch the surface of what R can do. Don’t be scared to try some coding of your own: you can always hit CTRL+Z to undo your edits!

There are 10 steps below that get increasingly tricky as you work through them.


STEP 1: Downloading and installing the software

Here’s where you want to go to get the software.

R: https://cran.ma.imperial.ac.uk/
R Studio: https://rstudio.com

Create a folder somewhere easy to find on your computer’s hard drive. Download the example data file (prog-example-data.csv) from this repository and move it into that folder.

STEP 2: Understanding a few basic “need-to-knows”

Any text that is prefixed with a hash mark “#” will be ignored by R, so you can use them to annotate code.

R is CaSe SenSitVe so if you run into an error, check your spelling and/or capitalizations.

To run code, either:

  • Place the cursor into a single chunk of code you want to run and hit CTRL+RETURN.
  • Select (highlight) chunks of code that you want to run and click the “Run” button.
  • Place the cursor at the start of your code and click “Run“: it’ll run everything!

If at first you don’t succeed, copy and paste any error you receive into a search engine! Whatever you’ve done, someone else will have done it and explained how to fix it!

STEP 3: Familiarizing yourself with the typical R Studio layout

When you open R Studio, it should look a bit like this. But if it doesn’t, don’t worry, it soon will!

Let’s also make sure you’ve got a .R file open to write your code in. Either navigate to File > New File > R Script or press CTRL+SHIFT+N.

Now you should have four quadrants:

  • Top left: This is your .R text file, where you’ll write and save your code.
  • Bottom left: This is the console, where you’ll see you code executed and get your data outputs.
  • Top right: This is your environment, where your data and other objects you’ve created will be listed.
  • Bottom left: This is your file structure, where you’ll see the working directory you’re in (where R will look for data and write files to, by default). There are also tabs plots and in-software help.

STEP 4: Set a working directory

We need to tell R Studio where in our hard drive or online to go to find the things we want (code, data, etc.)

  • Navigate to the “Go to Directory” button on the “Files” tab of the bottom-right pane (the three dots).
  • From there, navigate to the folder that you created and into which you saved the data. Then click “Open“.
  • In the “Files” tab (bottom-right pane) click More > Set As Working Directory. You’ll see in the console (bottom-left pane) that this is executed like this:
> setwd(~R\Training) 

STEP 5: Loading some data into R Studio

Now we can load the data into our project environment so that we can run code to explore it. Type the following into you .R text file.

mydata <- read.csv("prog-example-data.csv", header = TRUE)  
View(mydata)  

Objects are created using “<-” or “=“. This code creates an object mydata, a dataframe representing the .csv file.

The header = T option let’s R know that your columns have headers at the top. The second line allows you to view the dataframe object that you’ve just created. You should see it in the environment (top-right pane).

Tip: The = operator alone doesn’t mean equal to in R, it means assign to. To identify something as equal to use “==“. Consider other operators that are combinations of two: not-equal-to: “!=“; less-than-or-equal-to: “<=“. Think of “==” as “equals is-equals-to“!

Base R also allows you to calculate and use simple descriptive information for a variable like mean, median, standard deviation, minimum, maximum, etc. The “$” operator separates an object from the elements within it: object$element. Here, that’s in the format dataframe$variable. Or, for many functions, you can call a summary function and get a range of descriptive information (if you’ll excuse the pun).

mean(mydata$age)
median(mydata$age)
max(mydata$age)
min(mydata$age)

my_mean <- mean(mydata$age)
my_sd <- sd(mydata$age)

one_sd_below <- my_mean - my_sd
print(one_sd_below)
one_sd_above <- my_mean + my_sd
print(one_sd_above)

summary(mydata$age)

STEP 6: We need packages and their functions

Packages are the equivalent of the menus of functions you select from the toolbar in SPSS. We’ll load and use the stats package. We can install the package to the “library” of packages that R stores on your hard drive, then load them from that library to use its functions in our code. Specifying “dependencies = TRUE” tells R that you also want to install any other packages that the stats package relies on to run.

install.packages("stats", dependencies = TRUE)
library(stats)  

STEP 7: Run a between-subjects t-test

Ok, let’s test the null hypothesis that there is no statistically-significant group association between antisocial personality disorder (apd) and whether or not the individual started an intervention (prog.type). For this t-test we’re testing for an association between a numerical score of whole numbers and a categorical variables, so first we want to check our data are numerical so that the test will execute correctly:

class(mydata$apd)
class(mydata$prog.type)

Hopefully, this has returned something like:

> class(mydata$apd)
[1] "integer"
> class(mydata$prog.type)
[1] "factor"

If it doesn’t look like that, executing the following code tells R to recreate the variable adp as an integer. You can use similar code to recreate the prog.type variable as a “factor”.

mydata$adp <- as.integer(mydata$adp)
mydata$prog.type <- as.factor(mydata$prog.type)  

To run the t-test, type and run the following code:

t.test(mydata$apd ~ mydata$prog.type)

This should return something like:

> t.test(mydata$apd ~ mydata$prog.type)

		Welch Two Sample t-test

data:  mydata$apd by mydata$prog.type
t = -0.3321, df = 663.07, p-value = 0.7399
alternative hypothesis: true difference in means is not equal to 0
95 percent confidence interval:
  -0.4276355  0.3039068
sample estimates:
mean in group izon mean in group None 
          4.543046           4.604911

This tells us that the group means are 4.5 for the programme group and 4.6 for the no-programme group. The t-value is -0.33 at 663.07 degrees of freedom, and the p-value is .740. We can also see that the 95% confidence intervals cross zero, which indicates that the true effect could be either negative or positive: [-0.43, 0.30]. So, we fail to reject the null hypothesis.

Feel free to try a different t-test by y selecting different numerical and categorical variables in the dataset. You could also try a within-subjects t-test using a comma instead of a tilde (~) and specifying a paired test: for example, pre-and post-programme difference in problem solving scores (pre.prob & post.prob):

t.test(mydata$pre.prob, mydata$post.prob, paired = TRUE)

Analysis of variance (ANOVA), for where you have a categorical variable with more than two groups, is a bit more tricky – as is the case with SPSS. You can run a simple ANOVA, using the example code below, to get an omnibus F value and corresponding p-value. But you would want to specify your post-hoc contrasts etc. in a comprehensive ANOVA. Not difficult in R, but for another time and place.

my_aov <- aov(mydata$apd ~ mydata$risk.gen) # you can also do anova for more than one groups
summary(apt_aov)

STEP 8: Let’s try a basic chi-square test of association!

How about basic associations between different categorical variables? Let’s try the pre-treatment risk category (risk.gen) and participation in an intervention (prog.type). We”ve establishedm that R recognises prog.type as categorical so let’s check the risk variable is also in the correct format:

class(mydata$risk.gen)

This should also return that risk.gen is a “factor”. However, R may have identified the labels as text and decided to classify it is a “character” variable. If so, we can convert it into a factor like this:

mydata$risk.gen <- as.factor(mydata$risk.gen)

We can now examine what groups are in that factor with either a str (string) command or by examining the levels:

str(mydata$risk.gen)
levels(mydata$risk.gen)

The output tells us that risk.gen is a factor with four groups, but that those groups are in alphabetical order and not their natural order of magnitude:

> str(mydata$risk.gen)
 Factor w/ 4 levels "High","Low","Medium",..: 1 3 1 1 1 1 1 4 3 3 ...
> levels(mydata$risk.gen)
[1] "High"   "Low"    "Medium" "Vhigh"

We can fix that by specifying risk.gen as a factor and specifying the levels we expect the variable to have. Note that “c” means “combine”. We’re asking R to replace the current combination of four text labels with this new combination of four text labels.

mydata$risk.gen <- as.factor(mydata$risk.gen, levels = c("Low", "Medium", "High", "Very high")  

We can also look at a table of the data before we run the chi-square test:

table(mydata$risk.gen, mydata$prog.type)

The output shows us the frequencies within each level of the factor, by iZon programme participant or no-programme groups.

> table(mydata$risk.gen, mydata$prog.type)
    
         izon None
  High     81  108
  Low      11   26
  Medium  170  242
  Vhigh    40   72

Now we run a basic default chi-square test, along with an accompanying Fisher’s exact test, by executing:

chisq.test(mydata$risk.gen, mydata$prog.type)
fisher.test(mydata$risk.gen, mydata$prog.type)

Like the t-test, the output gives us a chi-square test value, degrees of freedom, and a p-value. The Fisher’s exact test returns a p-value.

> chisq.test(mydata$risk.gen, mydata$prog.type)

	Pearson's Chi-squared test

data:  mydata$risk.gen and mydata$prog.type
X-squared = 9.5829, df = 9, p-value = 0.3853

> fisher.test(mydata$prog.type, mydata$risk.gen)

	Fisher's Exact Test for Count Data

data:  mydata$prog.type and mydata$risk.gen
p-value = 0.3459
alternative hypothesis: two.sided

Once again, our p-value of 0.39 exceeds the p < 0.05 threshold and we fail to reject the null hypothesis. The Fisher’s exact test supports this decision with a p-value of 0.35. Feel free to try your own chi-square tests on other categorical variables (e.g., previous convictions).

STEP 9: We can use ggplot2 to chart those effects

One of the strengths of R is the ability to visualise data and outcomes. But the default plot function in R is not a good example of that strength. First we’ll run some code for a simple-but-ugly base chart, then we’ll run some for a complex-but-beautiful ggplot chart! But first we need to install the ggplot2 package:

install.packages("ggplot2", dependencies = TRUE)
library(ggplot2)  

Here’s the plot you can get from base R. It should arrive in a new tab in the bottom-right hand window of R Studio.

plot(mydata$prog.type, mydata$apd)

Probably couldn’t get that into a journal article right? We can spruce it up with ggplot2. We’ll also install and load ggpubr and Hmisc, so we can include fancy elements like significance bars and error bars.

install.packages(c("ggpubr", "Hmisc"), dependencies = TRUE)
library(ggpubr)  
library(Hmisc)  

You can still run quite basic plots with ggplot. Here’s the basic code for the adp t-test, specifying that you want the chart to plot the summary values derived from the mean function:

ggplot(mydata, aes(prog.type, apd) +
  geom_bar(stat = "summary", fun = "mean")

Also not as visually attractive as it could be. Let’s add some bells and whistles from ggplot2:

mycolours <- c("#284969", "#5599C6") # colours
prog.comps <- list(c("izon", "None")) # comparison for significance
ggplot(mydata, aes(prog.type, apd, fill = prog.type)) +
  stat_summary(fun = mean, geom = "bar", position = "dodge", alpha = 0.5, colour = "black") +
  stat_summary(fun.data = mean_cl_normal, geom = "errorbar", position = position_dodge(width = .90), width = .1) +
  stat_compare_means(comparisons = prog.comps, method = "t.test",
                     label = "p.signif", label.y = c(5.5), 
                     tip.length = .01, size = 4) +
  labs(x= "Prog", y = "APD", fill = "Prog") +
  scale_fill_manual(values = mycolours, labels = c("izon", "None")) +
  theme(text = element_text(size = 12), legend.position = "none")

This should arrive in the “Plots” tab in the bottom-right hand pane. This is a bit more complicated, so let’s have a quick look at the ingredients here:

  • Line 1: Creates a vector object called mycolours that specifies hash values for two colours.
  • Line 2: Creates a list object called prog.comps that will tell the ggpubr package how to render your significance bar.
  • Line 3: Starts the chart as an object. Specifies mydata as the source of the data, and the aesthetics (aes) as the variables prog.type and apd. Specifies the prog.type variable as the basis for grouping bars. The plus sign (+) at the end of each line tells R to join these bits of code into one “chunk”.
  • Line 4: Specifies the stats you want to the chart to use: chart the mean function from the stats package, as a bar chart. This includes the options to separate the bars (dodge) make them 50% transparent (alpha), and give the bars a black outline (colour).
  • Line 5: Adds the error bars, using the mean using the mean_cl_normal function from the Hmisc package, and instructions on where to position them and their size.
  • Line 6: Adds the significance bar by specifying this is the outcome of a t.test and that you want bars for each combination of the levels specified in prog.comps. It also specifies that you want the display the p-value. Lastly it specifies how high up the y-axis to put those bars (label.y) and what size and shape they should be (size and tip.length).
  • Line 7: Adds labels to the chart on the x- and y-axis and specifies what labels to use for the bars (fill).
  • Line 8: Adds the colours specified in the mycolours object and which levels you want them applied to.
  • Line 9: Lastly, changes the size of the text and removes the legend.

Dare to try changing some of the visualisations?! For example, try changing legend.position = "none" to legend.position = "top"…

STEP 10: Save your code file and consider your own journey into the brave new world…

Save your code file by navigating to File > Save As and save your file with a name like “training.R”. (Who knows, you might be coming back to add to it in the future!)

We’ve come to the end of this whistle-stop tour. The possibilities for analysis on R are vast, but the learning curve is incredibly steep. However, if you commit to integrating it into your stats projects and building on your skills, it really will start to come naturally – I promise!

Finally, I would recommend adding something like Andy Field’s excellent R version of his SPSS textbook to your library. It’s a really good investment. https://uk.sagepub.com/en-gb/eur/discovering-statistics-using-r/book236067


If you’ve got this far, congratulations and thank you! Good luck!