O'Reilly Book Coming Early 2021

Data Science on AWS

YouTube Videos, Meetups, Book, and Code: https://datascienceonaws.com

Workshop Description

In this workshop, we build a natural language processing (NLP) model to classify sample Twitter comments and customer-support emails using the state-of-the-art BERT model for language representation.

To build our BERT-based NLP model, we use the Amazon Customer Reviews Dataset which contains 150+ million customer reviews from Amazon.com for the 20 year period between 1995 and 2015. In particular, we train a classifier to predict the star_rating (1 is bad, 5 is good) from the review_body (free-form review text).

Workshop Cost

This workshop is FREE, but would otherwise cost <25 USD.

Workshop Agenda

Workshop Contributors

Workshop Instructions

1. Create `TeamRole` IAM Role

2. Launch an Amazon SageMaker Notebook Instance

Open the AWS Management Console

In the AWS Console search bar, type SageMaker and select Amazon SageMaker to open the service console.

In the Notebook instance name text box, enter workshop.

Choose ml.t3.medium (or alternatively ml.t2.medium). We'll only be using this instance to launch jobs. The training job themselves will run either on a SageMaker managed cluster or an Amazon EKS cluster.

Volume size 250 - this is needed to explore datasets, build docker containers, and more. During training data is copied directly from Amazon S3 to the training cluster when using SageMaker. When using Amazon EKS, we'll setup a distributed file system that worker nodes will use to get access to training data.

In the IAM role box, select the default TeamRole.

You must select the default VPC, Subnet, and Security group as shown in the screenshow. Your values will likely be different. This is OK.

Keep the default settings for the other options not highlighted in red, and click Create notebook instance. On the Notebook instances section you should see the status change from Pending -> InService

While the notebook spins up, continue to work on the next section. We'll come back to the notebook when it's ready.

3. Update IAM Role Policy

Click on the notebook instance to see the instance details.

Click on the IAM role link and navigate to the IAM Management Console.

Click Attach Policies.

Select IAMFullAccess and click on Attach Policy.

Note: Reminder that you should allow access only to the resources that you need.

Confirm the Policies

4. Start the Jupyter notebook

Note: Proceed when the status of the notebook instance changes from Pending to InService.

5. Launch a new Terminal within the Jupyter notebook

Click File > New > [...scroll down...] Terminal to launch a terminal in your Jupyter instance.

6. Clone this GitHub Repo in the Terminal

cd ~/SageMaker && git clone https://github.com/data-science-on-aws/workshop

Within the Jupyter terminal, run the following:

cd ~/SageMaker && git clone https://github.com/data-science-on-aws/workshop

7. Navigate Back to Notebook View

8. Start the Workshop!

Navigate to 01_setup/ in your Jupyter notebook and start the workshop!

You may need to refresh your browser if you don't see the new workshop/ directory.

Name		Name	Last commit message	Last commit date
Latest commit History 2,248 Commits
01_setup		01_setup
02_usecases		02_usecases
03_automl		03_automl
04_ingest		04_ingest
05_explore		05_explore
06_prepare		06_prepare
07_train		07_train
08_optimize		08_optimize
09_deploy		09_deploy
10_pipeline		10_pipeline
11_stream		11_stream
12_security		12_security
img		img
wip		wip
.gitignore		.gitignore
README.md		README.md

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

O'Reilly Book Coming Early 2021

Data Science on AWS

Workshop Description

Workshop Cost

Workshop Agenda

Workshop Contributors

Workshop Instructions

1. Create `TeamRole` IAM Role

2. Launch an Amazon SageMaker Notebook Instance

3. Update IAM Role Policy

4. Start the Jupyter notebook

5. Launch a new Terminal within the Jupyter notebook

6. Clone this GitHub Repo in the Terminal

7. Navigate Back to Notebook View

8. Start the Workshop!

About

Releases

Packages

Languages

License

anshuhan/data-science-on-aws

Folders and files

Latest commit

History

Repository files navigation

O'Reilly Book Coming Early 2021

Data Science on AWS

Workshop Description

Workshop Cost

Workshop Agenda

Workshop Contributors

Workshop Instructions

1. Create TeamRole IAM Role

2. Launch an Amazon SageMaker Notebook Instance

3. Update IAM Role Policy

4. Start the Jupyter notebook

5. Launch a new Terminal within the Jupyter notebook

6. Clone this GitHub Repo in the Terminal

7. Navigate Back to Notebook View

8. Start the Workshop!

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

1. Create `TeamRole` IAM Role

Packages