Machine Learning with R Tutorial: Introduction to k-means Clustering

4,850 views · Published 16 March 2017 · 2:29 · Indexed 10 October 2026

Channel: DataCamp · 2017 · Education

Watch on YouTube

Make sure to like & comment if you liked this video!

This is the second video for our course Unsupervised Learning in R by Hank Roark. Take Hank's course here: https://www.datacamp.com/courses/unsupervised-learning-in-r

Many times in machine learning, the goal is to find patterns in data without trying to make predictions. This is called unsupervised learning. One common use case of unsupervised learning is grouping consumers based on demographics and purchasing history to deploy targeted marketing campaigns. Another example is wanting to describe the unmeasured factors that most influence crime differences between cities. This course provides a basic introduction to clustering and dimensionality reduction in R from a machine learning perspective, so that you can get from data to insights as quickly as possible.

Now that we have some conceptual understanding of unsupervised learning and the different goals of unsupervised learning, let's dig right in with one popular approach to unsupervised learning.

K-means is a clustering algorithm, an algorithm used to find homogenous subgroups within a population. K-means is the first of two clustering algorithms to be covered in this course.

The K-means algorithm works by first assuming the number of subgroups, or clusters, in the data and then assigns each observation to one of those subgroups. In the next video, we will go deeper into how the k-means algorithm works to achieve this goal.

For example, one might hypothesize that this data shown on the screen contain 2 subgroups.

The k-means algorithm would assign all points in the top right hand corner to one subgroup and all observations in the bottom left hand corner to the other subgroup.

k-means in R comes with the base R install. Invoking k-means in R is simply a function call to ‘kmeans’ function, typically with three parameters.

The first parameter is the data, represented as ‘x’ here. In k-means, like many machine learning algorithms, the data is structured in a matrix with one observation per row of the matrix and one feature in each column of the matrix.

The next parameters for ‘kmeans’ is the number of predetermined groups or clusters. This parameter is called ‘centers’, for reasons that will be covered in the next video.

Finally, the kmeans algorithm has a random component. The implication of this stochastic component is that a single run of kmeans may not find the optimal solution to kmeans.

To overcome the random component of the algorithm, ‘kmeans’ can be run multiple times with the ‘best’ outcome across all runs being selected as the single outcome. ‘nstart’ is the parameter that specifies the number of times ‘kmeans’ will be repeated.

There are other parameters to ‘kmeans’ and I encourage you to check those out in the R documentation when you are ready.

The first exercises use synthetic data that were generated from three subgroups. But if you plot the data it might only appear to be two subgroups. Later in this chapter, you will see how k-means can be used to estimate the number of subgroups when the number of subgroups is not known a priori.

Later in this first chapter of the course, you will get experience applying ‘kmeans’ with a real world, but fun, dataset.

With that information, let's get started on the first exercise using ‘kmeans’.

More from this channel