diff --git a/README.md b/README.md index 75a90db..6d478d4 100644 --- a/README.md +++ b/README.md @@ -42,6 +42,7 @@ The first workshop supports multiple programming languages. Choose one: | **JavaScript** | Node.js 18+ | | **Python** | Python 3.10+ | | **C#** | .NET 8+ | +| **R** | R 4.2+ (see [R Data Analysis lab](labs/r-data-analysis.md)) | ## 📚 Workshop Structure @@ -117,6 +118,15 @@ Encode deep domain expertise: - Including resources (scripts, templates, data) alongside skills - `/skills` command and visibility controls +### R Track: Data Analysis + +### [Lab R - R Data Analysis](labs/r-data-analysis.md) +A standalone lab for R developers covering the same Copilot features through a data-analysis lens: +- Inline completion for dplyr pipelines and ggplot2 charts +- Copilot Chat for statistical summaries and code explanation +- Agent Mode to scaffold complete analysis scripts +- Next Edit Suggestions for repetitive transformations + ## 📁 Project Structure ``` @@ -133,7 +143,8 @@ copilot-katas/ │ ├── 06-instructions.md │ ├── 07-prompts.md │ ├── 08-agents.md -│ └── 09-skills.md +│ ├── 09-skills.md +│ └── r-data-analysis.md ← R track (standalone) ├── starter-code/ │ ├── javascript/ ← Todo App (Labs 00-04) │ │ ├── package.json @@ -145,6 +156,10 @@ copilot-katas/ │ ├── csharp/ ← Todo App (Labs 00-04) │ │ ├── TodoApp.csproj │ │ └── src/ +│ ├── r/ ← R Data Analysis (Lab R) +│ │ ├── README.md +│ │ ├── data/ +│ │ └── src/ │ └── recipe-api/ ← Recipe API (Labs 05-09) │ ├── package.json │ ├── README.md diff --git a/labs/r-data-analysis.md b/labs/r-data-analysis.md new file mode 100644 index 0000000..7974f27 --- /dev/null +++ b/labs/r-data-analysis.md @@ -0,0 +1,456 @@ +# Lab R – R Data Analysis with GitHub Copilot + +## Learning Goals + +By the end of this lab, you will be able to: +- Use inline completion to write idiomatic R data-wrangling code +- Use Copilot Chat to explore, clean, and transform a real dataset +- Generate statistical summaries and visualisations with Copilot's help +- Apply Agent Mode to build a complete analytical script end-to-end +- Write better R-specific prompts for more accurate Copilot suggestions + +## Introduction + +This lab is designed for developers who primarily work in **R**. Instead of a +generic Todo application, you will analyse a realistic sales dataset using the +tools and idioms R developers use every day: + +- `readr` / base `read.csv` for data loading +- `dplyr` for data manipulation +- `ggplot2` for visualisation +- `lubridate` for date handling +- Base R statistics (`mean`, `sd`, `cor`, `lm`, …) + +All exercises follow the same Copilot workflow as the other labs – but every +example is grounded in data analysis rather than application development. + +> **Note:** You do **not** need to complete the other labs before this one. +> This lab stands alone and covers inline completion, Chat, and Agent Mode in +> the context of R. + +## Prerequisites + +1. R 4.2 or later ([CRAN](https://cran.r-project.org/)) +2. VS Code with the + [R extension](https://marketplace.visualstudio.com/items?itemName=REditorSupport.r) +3. GitHub Copilot Chat extension installed and activated +4. Required packages – open an R console in your terminal and install them: + +```bash +# Start the R console +R + +# Then inside the R console, run: +install.packages(c("dplyr", "ggplot2", "lubridate", "readr", "forecast")) + +# When done, quit the R console +q() +``` + +> **Tip:** Type `R` in your terminal to start an interactive R session, and +> `q()` (or press `Ctrl+D`) to exit it. + +## Setting Up + +```bash +cd starter-code/r +``` + +Open `src/analysis.R` as your primary working file. This is the file +Copilot will use for most of its context. + +--- + +## Exercise 1: Loading and Exploring Data (Inline Completion) + +### Task 1.1: Load the Dataset + +Open `src/analysis.R` and let Copilot help you load the CSV file. + +Type the following comment and press **Enter**: + +```r +# Load the sales data from data/sample_sales.csv into a data frame called sales +``` + +**✅ Try This:** +1. Type the comment and press `Enter` +2. Wait for Copilot's inline suggestion (ghost text) +3. Press `Tab` to accept, or `Esc` to dismiss and try a different approach +4. Observe whether Copilot chooses `read.csv()` or `readr::read_csv()` – try + pressing `Alt+]` (Windows/Linux) / `Option+]` (Mac) to see alternative suggestions + +### Task 1.2: Explore the Data Structure + +After loading, add these comments one by one and let Copilot complete each: + +```r +# Show the first 6 rows of the sales data frame +``` + +```r +# Print a compact summary of the structure of sales (types, sample values) +``` + +```r +# Display descriptive statistics for all numeric columns +``` + +```r +# Count the number of rows and columns in the data frame +``` + +**💡 Tip:** Copilot reads the name `sales` as context. If you rename the +variable, Copilot will automatically use the new name in its suggestions. + +### Task 1.3: Unique Values and Distributions + +```r +# Print all unique values in the region column +``` + +```r +# Count how many transactions belong to each product category +``` + +```r +# Find the date range covered by this dataset (earliest and latest date) +``` + +**🔬 Experiment:** Try asking for the same thing in different ways: +- `# List unique regions` vs `# Get distinct values from the region column` + +Notice how the phrasing affects the suggested function (`unique()` vs +`dplyr::distinct()`). + +--- + +## Exercise 2: Data Cleaning and Transformation (Inline Completion) + +Open `src/data_processor.R`. You will implement several data-processing +functions here. + +### Task 2.1: Parse Dates + +The `date` column is loaded as a character string. Write a comment and let +Copilot convert it: + +```r +# Convert the date column from character to Date type using as.Date +``` + +**Observe:** Copilot should suggest the correct format string (`"%Y-%m-%d"`) +because it can see the sample data nearby. + +### Task 2.2: Add a Revenue Column + +Revenue = `units_sold × unit_price × (1 – discount)`. + +```r +# Add a revenue column: units_sold * unit_price * (1 - discount) +``` + +Try making the comment more specific and see how the suggestion changes: + +```r +# Add a new column called 'revenue' that equals units_sold multiplied by unit_price multiplied by (1 minus discount), using dplyr mutate +``` + +### Task 2.3: Handle Missing Values + +```r +# Remove rows that contain any NA values from the data frame +``` + +```r +# Print the number of rows before and after removing NAs +``` + +### Task 2.4: Implement the clean_sales_data Function + +Now wrap all the transformations into a reusable function. Type: + +```r +# Function to clean the raw sales data frame: +# - Convert date column to Date type +# - Add a revenue column +# - Remove rows with NA values +# Returns the cleaned data frame +clean_sales_data <- function(df) { +``` + +Press `Enter` inside the function body and let Copilot suggest the +implementation. Use **partial accept** (`Ctrl+→` on Windows / `Cmd+→` on Mac) +to accept one expression at a time. + +**💡 Tip – partial accept:** + +| Shortcut | Action | +|---|---| +| `Cmd+→` / `Ctrl+→` | Accept next word | +| `Tab` | Accept entire suggestion | +| `Esc` | Dismiss and keep typing | + +--- + +## Exercise 3: Statistical Summaries (Chat – Ask Mode) + +Open Copilot Chat (`Ctrl+Shift+I` / `Cmd+Shift+I`) and make sure **Ask Mode** +is selected. + +### Task 3.1: Ask About Your Data + +With `src/analysis.R` open (and mentioned as context), try these prompts: + +``` +What dplyr pipeline would calculate total revenue grouped by region, +sorted from highest to lowest? +``` + +``` +How do I calculate the mean, median, and standard deviation of units_sold +for each product category? +``` + +``` +Show me how to find the top 3 sales representatives by total revenue +``` + +**✅ Observe:** +- How Chat structures multi-step dplyr pipelines +- How it adapts suggestions to your column names + +### Task 3.2: Use Slash Commands + +**`/explain`** – Select one of the pipelines Chat suggested and type: +``` +/explain +``` + +**`/fix`** – Introduce a deliberate bug: +```r +# Bug: wrong column name +total_by_region <- sales |> + group_by(area) |> # 'area' does not exist + summarise(revenue = sum(revenue)) +``` +Select the code and type `/fix`. + +**`/doc`** – Select `clean_sales_data` and type: +``` +/doc +``` +Observe how Copilot generates roxygen2-style documentation. + +**`/tests`** – Select a function and type: +``` +/tests using testthat +``` + +### Task 3.3: Context References + +Use `#file:` to reference a specific file as context in Copilot Chat. Start +typing `#` and select the file from the autocomplete list. + +``` +#file:data_processor.R Does this file handle the case where the +discount column contains values outside the range 0–1? +``` + +``` +@workspace What functions have been defined so far across all R files? +``` + +--- + +## Exercise 4: Data Visualization (Inline Completion) + +Open `src/visualizer.R`. + +### Task 4.1: Monthly Revenue Trend + +```r +# Plot monthly total revenue as a line chart using ggplot2 +# x-axis: month (extract from date column), y-axis: total revenue +``` + +Copilot should suggest something involving `lubridate::floor_date` or +`format(date, "%Y-%m")` to aggregate by month. + +### Task 4.2: Revenue by Region + +```r +# Create a bar chart of total revenue by region, sorted from highest to lowest +# Use reorder() to sort the bars +``` + +### Task 4.3: Top Products + +```r +# Horizontal bar chart of the top products by total revenue +# Flip coordinates so product names are on the y-axis +``` + +> **Note:** The sample dataset contains only 4 unique products, so all +> products will appear in the chart. + +### Task 4.4: NES – Pattern Recognition + +1. Create two ggplot charts with similar structure +2. Edit the title of the first chart +3. Look for the **Next Edit Suggestion** indicator – Copilot may suggest + updating the second chart's title in the same way +4. Press `Tab` to accept + +--- + +## Exercise 5: Build a Complete Analysis Script (Agent Mode) + +Switch Copilot Chat to **Agent Mode** for this exercise. + +### Task 5.1: Generate a Full Report Script + +In Agent Mode, paste this prompt: + +``` +I have a sales CSV at data/sample_sales.csv with columns: +date, region, product, category, units_sold, unit_price, discount, sales_rep. + +Create a complete R script called src/report.R that: +1. Loads the data using readr::read_csv +2. Cleans it (parse dates, add revenue column, drop NAs) +3. Prints a summary table: total revenue and units sold per region +4. Prints the top products by revenue +5. Creates the output/ directory if it doesn't exist +6. Saves a ggplot2 bar chart of revenue by region as output/revenue_by_region.png +7. Prints a simple linear regression of revenue ~ units_sold and shows the R² + +Use dplyr and ggplot2. Add comments explaining each step. +``` + +> **Note:** You may notice a very low R² for the `revenue ~ units_sold` model. +> This is expected – revenue also depends on `unit_price`, which varies widely +> across products. Try asking Copilot Chat why the R² is low and how to improve +> the model (e.g., by adding `unit_price` as a predictor). + +**Observe Agent Mode:** +- Copilot creates the file autonomously +- It may ask clarifying questions before proceeding +- It runs the file and iterates if there are errors + +### Task 5.2: Iterative Refinement + +After reviewing the generated script, ask Copilot to improve it: + +``` +Add a second chart showing monthly revenue trend as a line plot +and save it as output/monthly_trend.png +``` + +``` +Wrap the cleaning and aggregation steps in named functions and +move them to data_processor.R +``` + +--- + +## Exercise 6: Advanced Techniques + +### Task 6.1: Inline Chat for Quick Edits + +Press `Ctrl+I` / `Cmd+I` to open inline chat on a selected code block. + +**Try these prompts on your summary pipeline:** +- `"Make this pipeline handle the case where revenue is NA"` +- `"Rewrite this using base R instead of dplyr"` +- `"Add a percentage-of-total column to this summary"` + +### Task 6.2: Specificity Experiment + +Compare how vague vs. specific comments affect suggestions: + +**Vague:** +```r +# filter the data +``` + +**Specific:** +```r +# Filter to keep only Electronics category rows where revenue > 1000 +``` + +**Very specific:** +```r +# Filter the sales data frame to rows where category equals "Electronics" +# and revenue is greater than 1000, returning a new data frame +``` + +### Task 6.3: Statistical Modelling + +Let Copilot help you build a simple model. Type these comments one at a time: + +```r +# Fit a linear regression model to predict revenue from units_sold +``` + +```r +# Print the model summary including R-squared and coefficient p-values +``` + +```r +# Create a residuals vs fitted plot to check model assumptions +``` + +--- + +## Challenges + +### Challenge 1: Sales Rep Performance Dashboard + +Write a comment block describing a function that takes the cleaned sales data +and returns a data frame ranking each sales rep by: +- Total revenue +- Average discount given +- Number of transactions +- Most sold product category + +Let Copilot generate the full implementation. + +### Challenge 2: Outlier Detection + +Ask Copilot Chat: + +``` +How would I detect and flag outlier transactions in this dataset +using the IQR method on the revenue column? +``` + +Implement the suggestion and test it with the sample data. + +### Challenge 3: Time-Series Forecasting + +> **Note:** This challenge requires the `forecast` package which was included +> in the prerequisites install command. If you skipped it, run: +> `install.packages("forecast")` first. + +``` +Using the monthly revenue data I have, show me how to fit a simple +exponential smoothing model with the forecast package and plot +a 3-month ahead forecast +``` + +--- + +## Key Takeaways + +| Concept | R-specific tip | +|---|---| +| **Name your variables clearly** | `clean_sales` vs `df2` – Copilot uses variable names as context | +| **Comment-driven development** | Describe the transformation in plain English before writing code | +| **Specify packages** | Mention `dplyr`, `ggplot2`, `data.table` in comments to steer suggestions | +| **Partial accept** | Accept one pipe stage at a time with `Ctrl+→` | +| **Chat for pipelines** | Long dplyr/ggplot2 chains are ideal for Chat; iterate conversationally | +| **Agent Mode for scripts** | Let Agent Mode scaffold entire analysis scripts, then refine | + +--- + diff --git a/starter-code/r/README.md b/starter-code/r/README.md new file mode 100644 index 0000000..0c175b1 --- /dev/null +++ b/starter-code/r/README.md @@ -0,0 +1,68 @@ +# GitHub Copilot Katas – R Data Analysis + +A self-contained R project for the **R Data Analysis** lab exercise. + +## Prerequisites + +- R 4.2 or later ([Download](https://cran.r-project.org/)) +- VS Code with the [R extension](https://marketplace.visualstudio.com/items?itemName=REditorSupport.r) + +## Installing Packages + +Open an R console in your terminal and install the required packages: + +```bash +# Start the R console +R +``` + +```r +# Inside the R console, run: +install.packages(c("dplyr", "ggplot2", "lubridate", "readr", "forecast")) + +# When done, quit the R console +q() +``` + +> **Tip:** Type `R` in your terminal to start an interactive R session, and +> `q()` (or press `Ctrl+D`) to exit it. + +## Running the Analysis + +From the terminal (from this directory): + +```bash +Rscript src/analysis.R +``` + +Or open `src/analysis.R` in VS Code and run sections interactively using the +R extension. + +## Project Structure + +``` +r/ +├── README.md +├── data/ +│ └── sample_sales.csv ← provided sample dataset +└── src/ + ├── analysis.R ← main script (start here) + ├── data_processor.R ← data cleaning & transformation functions + └── visualizer.R ← ggplot2 visualisation functions +``` + +## Dataset + +`data/sample_sales.csv` contains fictional quarterly sales records with the +following columns: + +| Column | Description | +|--------------|------------------------------------| +| date | Transaction date (YYYY-MM-DD) | +| region | Sales region (North/South/East/West)| +| product | Product name | +| category | Product category | +| units_sold | Number of units sold | +| unit_price | Price per unit (USD) | +| discount | Discount applied (0.0 – 1.0) | +| sales_rep | Name of the sales representative | diff --git a/starter-code/r/data/sample_sales.csv b/starter-code/r/data/sample_sales.csv new file mode 100644 index 0000000..2fb2308 --- /dev/null +++ b/starter-code/r/data/sample_sales.csv @@ -0,0 +1,41 @@ +date,region,product,category,units_sold,unit_price,discount,sales_rep +2024-01-03,North,Widget A,Electronics,12,149.99,0.0,Alice +2024-01-05,South,Gadget B,Electronics,8,89.99,0.05,Bob +2024-01-07,East,Gizmo C,Hardware,25,34.99,0.0,Carol +2024-01-10,West,Widget A,Electronics,5,149.99,0.10,Dave +2024-01-12,North,Gadget B,Electronics,15,89.99,0.0,Alice +2024-01-14,South,Gizmo C,Hardware,30,34.99,0.0,Bob +2024-01-17,East,Widget A,Electronics,7,149.99,0.05,Carol +2024-01-19,West,Doohickey D,Tools,20,19.99,0.0,Dave +2024-01-21,North,Gizmo C,Hardware,18,34.99,0.0,Alice +2024-01-24,South,Widget A,Electronics,10,149.99,0.0,Bob +2024-02-01,North,Gadget B,Electronics,22,89.99,0.10,Alice +2024-02-04,South,Doohickey D,Tools,35,19.99,0.0,Bob +2024-02-06,East,Widget A,Electronics,9,149.99,0.0,Carol +2024-02-08,West,Gizmo C,Hardware,14,34.99,0.05,Dave +2024-02-11,North,Doohickey D,Tools,40,19.99,0.0,Alice +2024-02-13,South,Widget A,Electronics,6,149.99,0.0,Bob +2024-02-15,East,Gadget B,Electronics,11,89.99,0.0,Carol +2024-02-18,West,Widget A,Electronics,4,149.99,0.10,Dave +2024-02-20,North,Gizmo C,Hardware,28,34.99,0.0,Alice +2024-02-22,South,Doohickey D,Tools,50,19.99,0.05,Bob +2024-03-01,East,Widget A,Electronics,13,149.99,0.0,Carol +2024-03-03,West,Gadget B,Electronics,17,89.99,0.0,Dave +2024-03-05,North,Gizmo C,Hardware,22,34.99,0.0,Alice +2024-03-07,South,Widget A,Electronics,8,149.99,0.05,Bob +2024-03-10,East,Doohickey D,Tools,45,19.99,0.0,Carol +2024-03-12,West,Gadget B,Electronics,19,89.99,0.0,Dave +2024-03-14,North,Widget A,Electronics,11,149.99,0.0,Alice +2024-03-17,South,Gizmo C,Hardware,32,34.99,0.10,Bob +2024-03-19,East,Doohickey D,Tools,38,19.99,0.0,Carol +2024-03-21,West,Widget A,Electronics,7,149.99,0.0,Dave +2024-04-02,North,Gadget B,Electronics,24,89.99,0.0,Alice +2024-04-04,South,Widget A,Electronics,14,149.99,0.05,Bob +2024-04-06,East,Gizmo C,Hardware,29,34.99,0.0,Carol +2024-04-09,West,Doohickey D,Tools,55,19.99,0.0,Dave +2024-04-11,North,Widget A,Electronics,16,149.99,0.0,Alice +2024-04-13,South,Gadget B,Electronics,20,89.99,0.10,Bob +2024-04-16,East,Doohickey D,Tools,42,19.99,0.0,Carol +2024-04-18,West,Gizmo C,Hardware,17,34.99,0.0,Dave +2024-04-20,North,Widget A,Electronics,9,149.99,0.05,Alice +2024-04-23,South,Gizmo C,Hardware,26,34.99,0.0,Bob diff --git a/starter-code/r/src/analysis.R b/starter-code/r/src/analysis.R new file mode 100644 index 0000000..02221ee --- /dev/null +++ b/starter-code/r/src/analysis.R @@ -0,0 +1,57 @@ +# ============================================================ +# Sales Data Analysis +# +# This is the main entry point for the R data analysis exercises. +# +# GitHub Copilot Katas - R Version +# +# To run: +# Rscript src/analysis.R +# +# Exercises: +# 1. Load and explore sales data +# 2. Clean and transform the data +# 3. Compute sales statistics by region and product +# 4. Visualize trends (if ggplot2 is available) +# ============================================================ + +# Load required packages +# Hint: use library() - let Copilot suggest which packages to load + +# ============================================================ +# Exercise 1: Load and Explore the Data +# ============================================================ + +# Load the CSV file from data/sample_sales.csv into a data frame called `sales` +# Then let Copilot help you explore it (head, str, summary, nrow, ncol) + +# ============================================================ +# Exercise 2: Data Cleaning and Transformation +# ============================================================ + +# Source the data processing functions +source("src/data_processor.R") + +# Call the cleaning function and store the result +# Then verify that the date column is of type Date and +# that the revenue column has been added + +# ============================================================ +# Exercise 3: Aggregation and Summary Statistics +# ============================================================ + +# Use dplyr (or base R) to: +# - Calculate total revenue by region +# - Calculate total revenue by product category +# - Identify the top-performing sales representative +# - Find the best-selling product by units sold + +# ============================================================ +# Exercise 4: Visualization (optional – requires ggplot2) +# ============================================================ + +# Source the visualisation functions +source("src/visualizer.R") + +# Plot monthly revenue trend +# Plot revenue breakdown by region diff --git a/starter-code/r/src/data_processor.R b/starter-code/r/src/data_processor.R new file mode 100644 index 0000000..7a5ac00 --- /dev/null +++ b/starter-code/r/src/data_processor.R @@ -0,0 +1,32 @@ +# ============================================================ +# Data Processor +# +# This file will contain functions for cleaning and +# transforming the sales data. +# +# GitHub Copilot Katas - R Version +# +# Functions to implement (let Copilot help you!): +# - clean_sales_data(df) : parse dates, add revenue column, remove NAs +# - add_revenue_column(df) : calculate revenue = units_sold * unit_price * (1 - discount) +# - filter_by_region(df, region) : return rows matching the given region +# - summarise_by_category(df) : total revenue and units per category +# ============================================================ + +# ============================================================ +# Exercise 2: Implement the data processing functions below. +# Start by writing a comment describing each function, +# then let Copilot suggest the implementation. +# ============================================================ + +# Function to clean the raw sales data frame: +# - Convert the 'date' column to Date type +# - Add a 'revenue' column: units_sold * unit_price * (1 - discount) +# - Remove any rows that contain NA values + +# Function to add a revenue column to a sales data frame +# Revenue = units_sold * unit_price * (1 - discount) + +# Function to filter sales data by a specific region (case-insensitive) + +# Function to summarise total revenue and total units sold, grouped by category diff --git a/starter-code/r/src/visualizer.R b/starter-code/r/src/visualizer.R new file mode 100644 index 0000000..2c5cfac --- /dev/null +++ b/starter-code/r/src/visualizer.R @@ -0,0 +1,28 @@ +# ============================================================ +# Visualizer +# +# This file will contain functions for creating charts and +# plots from the sales data using ggplot2. +# +# GitHub Copilot Katas - R Version +# +# Functions to implement (let Copilot help you!): +# - plot_monthly_revenue(df) : line chart of total revenue per month +# - plot_revenue_by_region(df) : bar chart of revenue per region +# - plot_top_products(df, n) : bar chart of the top n products by revenue +# ============================================================ + +# ============================================================ +# Exercise 4: Implement the visualisation functions below. +# Write a comment describing the chart, then let Copilot +# suggest the ggplot2 implementation. +# ============================================================ + +# Function to plot monthly revenue as a line chart +# Hint: you may need to extract the month from the date column + +# Function to create a bar chart of total revenue by region, +# with bars sorted from highest to lowest + +# Function to create a horizontal bar chart of the top n products +# (default n = 5) ranked by total revenue