Anything you notice during exploration needs confirming on data you did not explore. The more you look, the more flukes you find.
ExampleFifty charts from an agent in a minute: some will show patterns that are pure chance.
Dictionary
100 terms · examples drawn from the course
100 shown
Anything you notice during exploration needs confirming on data you did not explore. The more you look, the more flukes you find.
ExampleFifty charts from an agent in a minute: some will show patterns that are pure chance.
The share of predictions that were correct. Misleading when one answer is very common.
ExampleA car-park model that always says 'full' is 96% accurate and never finds an empty space.
Also: summary statistic
Reducing many values to one summary, such as a mean, sum, count, minimum or maximum.
Exampledf['left'].mean() is the share of customers who left: about 26.5%.
Also: agent, LLM agent
An LLM that plans a step, uses a tool (runs code, queries data, calls an API), reads the result and decides what to do next, in a loop.
ExampleEarth Agent proposes the search space and objective, runs experiments on Earth through its API, reads the results and suggests the next experiment.
Also: AI
Named in 1956 as the umbrella for making machines act intelligently (rules, search, logic, learning). Today it usually means deep learning and generative AI.
ExampleWhen someone says 'AI', ask which circle they mean.
Also: axis=0, axis=1
The direction of an aggregation on a table. axis=0 runs down the rows and gives one answer per column; axis=1 runs across and gives one answer per row.
Exampleservices.sum(axis=1) counts how many services each customer has.
Also: no-brainer model, benchmark
The simplest sensible answer, used as the bar any real model must beat. A model is never good on its own, only better than something.
ExampleAlways predicting 'stays' is 73.5% accurate on our data and finds no leavers. In the smoothie story, always guessing the average is off by 85 calories.
Also: data bias
When historical data reflects unfair or unrepresentative decisions, a model learns them and applies them at scale.
ExampleA model trained on ten years of biased hiring decisions repeats the bias a thousand times a day.
Also: bands, buckets, cut, qcut
Turning a number into bands. cut uses edges you choose; qcut makes bands with equal numbers of rows.
ExampleTenure bands 0-12, 13-24, 25-48 and 49-72 months make the churn pattern readable at a glance.
Also: mask, filter
An array or column of True/False values used to select rows. The mean of a mask is the share of True values.
Examplehigh = charges > 80 selects high payers; churned[high].mean() is the share of them who left.
Also: segment check
Testing whether a pattern holds inside each segment, not only on average. Sometimes an overall pattern disappears or reverses within groups.
ExampleMonthly contracts leave more within every internet service type.
Also: imbalance
How evenly the labels are spread. With imbalance, a model can look accurate by always predicting the common label.
ExampleAbout 1 in 4 customers left; 96 of 100 car-park spaces are full.
A supervised task where the answer is a category.
ExampleSpam or not spam; will leave or will stay.
Also: k-means
Grouping similar rows together. k-means places k centres and assigns each row to the nearest one, repeating until stable.
ExampleSegmenting customers before anyone has defined the segments.
A new user or item with no history for a recommender to learn from.
ExampleWhat do you recommend to someone who signed up a minute ago?
Also: project steps
The seven steps of a simple ML project: frame the use, build the dataset, explore, design the evaluation, build the baseline, diagnose and improve, confirm and hand over.
ExampleSession 1 covers the first three steps; Sessions 2 and 3 cover the rest.
Also: contingency table, cross tabulation
A table of counts for two columns at once. With normalize='index', each row shows shares that add up to 1.
Examplepd.crosstab(df['PaymentMethod'], df['Churn'], normalize='index'): electronic check payers leave far more often.
Complex models need very large datasets and compute; organisations without them should keep models simple.
ExampleGoogle can afford a deep network; a 5,000-row customer table usually can't.
Also: pandas DataFrame, df
pandas' table: a set of named columns (Series) that share one row index.
Exampledf = pd.read_csv('telco_churn_teaching.csv') gives a DataFrame of 7,071 rows and 22 columns.
The line or surface a classifier draws between labels. Different algorithms are allowed different boundary shapes.
ExampleLinear model: one straight line. Decision tree: vertical and horizontal cuts. Neural network: a flexible curve.
Also: use statement
One sentence that says who uses a prediction, when, for how many, and for what goal. If a bracket stays empty, the project isn't ready.
ExampleEvery Monday we score active customers and the retention team calls the 100 most likely to leave in the next 30 days, so that we keep more revenue than the calls cost.
Also: DecisionTreeClassifier, tree
A model that learns a sequence of yes/no questions about the features. Without limits it can memorize the training data.
ExampleAn unlimited tree is 99.7% right on training customers and 72.6% on new ones; limited to depth 4 it gets about 79% on both.
Neural networks with many layers that learn their own features from raw data. Strong on images, sound and text; needs much data and compute.
ExamplePixels to edges to shapes to 'cat'. On a modest business table, tree ensembles usually match it.
Also: data mining
Exploring data to understand it and get inspired, without claiming the findings hold beyond the data at hand.
ExampleLooking at last quarter's sales by region.
Also: drift, data shift
When the data a model meets in use differs from the data it learned from. Performance degrades quietly, with no error message.
ExampleThe users a model learned from were mostly older; the users it now serves are mostly young.
Three questions asked at every project step: what can I do myself (level 1), what must I raise with a data scientist or the business (level 2), and what can an AI agent take on.
ExampleSplitting the data: I implement it; I raise whether 'new data' means future months or new customers; an agent can write the code.
Also: data type
The type of the values in an array or column, such as int64, float64 or text (object or str). One wrong value can change the type of a whole column.
ExampleTotalCharges is read as text because 11 rows contain a blank space instead of a number.
scikit-learn's ready-made baseline. With strategy='most_frequent' it always predicts the most common label.
ExampleDummyClassifier(strategy='most_frequent') scores 0.735 on the test set.
Also: duplicated, repeated records
Rows that repeat: exact copies, or the same key (such as a customer ID) appearing more than once, sometimes with conflicting values.
ExampleThe teaching data has 22 exact copies and 6 customers who appear twice with different monthly charges.
Also: fresh toast, spurious pattern
A metaphor for finding patterns that aren't there. People see faces in toast; flexible algorithms find patterns in noise. Only fresh data exposes them.
ExampleA model trained on coin-flip labels scores 100% on what it studied and 49% on new data.
Also: raise
Raising a consequential choice with someone who can decide it. Knowing when to escalate is a skill, not a weakness.
ExampleWhether a downgrade counts as leaving goes to the retention lead, not to the analyst alone.
Also: fit, predict, scikit-learn API
The four methods every scikit-learn model shares: learn from examples, label new rows, give a score per label, and report accuracy.
ExampleSwapping LogisticRegression for DecisionTreeClassifier changes one line; the rest of the code stays the same.
Also: exploration
Looking at the data to understand it: distributions, comparisons between groups, and checks across segments. Done on training data only.
ExampleShare who left by contract, by tenure band and by internet service, each with counts.
Also: input, variable, column, X
A piece of information about each instance that the model can use as input. It must be known at the moment of prediction.
Exampletenure, Contract and MonthlyCharges are features; LastContactReason is not, because it is recorded after the outcome.
Also: domain knowledge
Creating inputs that help the model, usually from domain knowledge. A new column is a hypothesis written in code.
ExampleThe smoothie: adding grams of fat, carbs and protein cut the error from 47 to 4 calories, and the model rediscovered nutrition labels.
Retraining a model's parameters on your own examples. Powerful and costly; it turns the work into a supervised learning project.
ExampleTeaching a model a consistent house style from thousands of approved answers.
Also: shape, head, describe, value_counts
Five commands to run on any new data before anything else: shape, head, dtypes (or info), describe and value_counts.
ExampleOn our data they reveal 7,071 rows, a money column stored as text, and that most customers are on monthly contracts.
Before believing any number: missing is not zero, the same thing twice, implausible values, categories that hide structure, identifiers posing as features.
ExampleIn the teaching data: 11 blank totals, 28 repeated rows, and a 'No internet service' value in six columns.
Also: one job
How well a model does on data it has never seen. Machine learning has one job: succeed on new data.
ExampleDay 61: for days 1 to 60 you look the dose up; a model is only useful if the pattern still holds on day 61.
Also: split-apply-combine
Split rows into groups, compute a summary for each group, and combine the results into one table.
Exampledf.groupby('InternetService')['left'].agg(['mean', 'size']): fibre about 42% left out of 3,113 customers.
When an LLM produces fluent, confident text that is false.
ExampleInventing a policy clause that isn't in the HR handbook.
A stated guess about why the outcome happens, written before looking at the data and linked to a column that could test it.
Example'Customers leave because they aren't locked in by a contract' is tested with the Contract column.
Also: ID, key
A column that names a row, such as a customer ID. It carries no meaning about the outcome, but a flexible model can memorize it.
ExamplecustomerID is kept as a key for joins and duplicate checks, never used as a feature.
Also: outliers, sanity checks
Values that can't be right: impossible ranges, wrong units, misread formats. A domain expert often spots them in seconds.
ExampleDates in the year 0040 that were really two-digit years; a negative age; a premium plan billed at 0.
Also: row labels
The row labels of a DataFrame or Series. Operations line rows up by index, not by position.
ExampleAfter filtering, the index keeps the original row numbers, which is why loc and iloc can give different rows.
Also: row, example, observation
One example the model learns from or makes a prediction for: one row of the dataset.
ExampleOne customer in our data; one email in a spam filter.
The running Python process behind a notebook. It holds everything in memory: variables, imported libraries, loaded data.
ExampleRestarting the kernel wipes memory, which is how you find cells that depend on something run earlier.
Also: target, ground truth, y
The answer the model learns to predict. Someone has to define it: what counts, over which window, from which moment.
ExampleChurn = 'left within the month before the snapshot' in our data; our project needs 'cancels within 30 days of the scoring Monday'.
Also: target definition
The precise rule for what counts as a positive example, including ambiguous cases. It is a business decision, not a fact found in the data.
ExampleDoes a downgrade count as leaving? A customer who cancels and returns two weeks later? Someone has to rule.
Also: noisy labels
Errors or inconsistencies in the recorded answers.
Example'Spam' is whatever users moved to the spam folder; bakery sales undercount demand on sold-out days.
Also: language model
A model trained to predict the next word over a vast library of text, then tuned on examples and human preferences. Fluent is not the same as correct.
ExampleAsk the same question twice and get two different answers; when it doesn't know, it still sounds sure.
Autonomous (the agent does it and reports), semi-autonomous (the agent proposes, a person approves), manual (a person decides). Autonomy follows reversibility.
ExampleTrying hyperparameters is autonomous; defining success is manual.
Also: selection, indexing
Two ways to select rows and columns: loc by label or condition, iloc by position.
Exampledf.loc[df['tenure'] == 0, ['customerID', 'TotalCharges']] selects brand-new customers and two columns.
Also: LogisticRegression
A simple, fast classification model that combines the features with learned weights and outputs a probability.
ExampleOn our split it reaches about 80% test accuracy, against 73.5% for the baseline.
Also: ML
A way of getting a recipe that turns inputs into outputs: instead of a person writing the rules, an algorithm stitches them together from examples with answers.
ExampleNobody writes rules for spam; the filter learns them from emails people marked as spam.
Also: map
Replacing values using a dictionary. Anything missing from the dictionary becomes NaN, which works as an alarm.
Exampledf['PaperlessBilling'].map({'Yes': 1, 'No': 0}) turns text into 1/0.
Also: chaining
Writing a sequence of pandas steps as one expression, one step per line, so the whole recipe reads top to bottom.
Examplepd.read_csv(path).assign(TotalCharges=...) builds the cleaned table in one readable block.
A blank can mean 'not yet', 'unknown' or 'not applicable'. Filling it with 0, the average, or dropping the row are three different claims about the world.
ExampleBlank TotalCharges belong to customers in their first month: set to 0 ('no bill yet'), never to the average.
Also: NaN, null, NA
A value that isn't there. pandas shows it as NaN. Missing data can also hide as a blank space, a zero or a placeholder word.
Exampleisna() finds nothing in TotalCharges because the blanks are spaces; converting to numbers reveals 11 missing values.
Also: recipe
The recipe that comes out of machine learning: instructions that turn inputs into outputs. Once learned, it is used like any other program.
ExampleThe fitted logistic regression is a model; given a new customer's features it outputs a score.
Layers of simple mathematical units connected by learned weights. Allows very flexible decision boundaries.
ExampleThe squiggly boundary in the wine example.
Also: Jupyter, Colab notebook
An interactive document that mixes code cells, their outputs and text. Cells can run in any order; the number in brackets shows the order they actually ran.
ExampleA variable defined in a lower cell works in an upper one only because it was run earlier. Restart and run all to check the notebook really works top to bottom.
Also: ndarray, array
A grid of values that all share one type, stored so that whole-array operations run in fast compiled code.
Examplechurned = (df['Churn'] == 'Yes').to_numpy() gives an array of True/False values, one per customer.
Also: dummy variables, get_dummies
Turning a category into several 0/1 columns, one per value, so a model can use it without inventing an order.
ExampleContract becomes Contract_One year and Contract_Two year; month-to-month is the row with both zeros.
Also: memorizing
When a model memorizes quirks of its training data instead of learning patterns that hold on new data. Great on what it studied, poor on anything else.
ExampleThe unlimited decision tree, or the student who learns that 8 divided by 4 is infinity by turning the 8 on its side.
Also: make_pipeline
Preparation steps and a model packaged as one object, so new data is treated exactly the same way as training data. Covered properly in Session 2.
Examplemake_pipeline(StandardScaler(), LogisticRegression()) scales and models in one fit.
Also: pivot_table
A two-way summary: rows by one column, columns by another, any aggregation in the cells.
ExampleShare who left by contract type within each internet service: the monthly-contract effect holds in every group.
Also: predict_proba, score
The model's estimate of how likely each label is. predict() turns it into a label using a cut-off of 0.5 unless told otherwise.
ExampleA month-to-month fibre customer scores 0.68 'likely to leave'; the retention team might call the top 100 scores whatever the cut-off.
Also: horizon, window
How far ahead the label looks from the prediction point.
Example'Cancels within the next 30 days' has a 30-day horizon.
Also: prediction time, scoring time
The moment a prediction is made. Only information available at that moment may be used as a feature.
Example'Every Monday': anything recorded after Monday, such as a cancellation call, is off limits.
Also: decision log
A written record of every data fix: what, how many rows, what was done, and why. Every fix is a decision.
Example'Blank TotalCharges, 10 rows, set to 0, all tenure 0: no bill yet.'
One line per cycle recording what was decided and why. It becomes the project file that later sessions build on.
Example'Test set locked: 20%, stratified, seed 42.'
Also: levels
Four levels of ML projects: routine application, applied DS decisions, specialist ML, research. The programme trains full capability at level 1 and foundations of level 2.
ExampleBuilding a churn model on agreed data is level 1; choosing how to define churn when the business disagrees is level 2.
Also: prompt engineering
Steering an LLM with instructions and examples in the request. Cheapest way to adapt it; can be fragile.
ExampleAdding three examples of good answers to the request.
Also: random_state, seed
A fixed starting number for a random generator, so that 'random' results come out the same every run.
Examplerandom_state=42 makes the train/test split identical each time, so results can be compared.
Also: small-sample caution
Always show how many rows sit behind a percentage. Small groups produce extreme rates by chance.
Example67% of 12 customers says almost nothing; 42% of 3,113 customers is a pattern.
Also: recommender
A model that predicts the missing cells of a user-by-item table and shows each user their best few. Content-based, collaborative or hybrid.
Example'Customers also bought'. Traps: cold start, popularity bias, feedback loops.
A supervised task where the answer is a number. (In statistics the word means fitting a line; in ML it just means a numeric answer.)
ExampleCalories in a smoothie; tomorrow's demand for a bakery item.
Also: RL
Learning by trial and error from a score for actions, often delayed, where actions change what the learner sees next.
ExampleA system learning the game Breakout discovers a tunnel strategy nobody taught it.
Also: sampling bias
Whether the training data looks like the world the model will be used in. The training data is the only world a model knows.
ExampleUsers from an island near Antarctica, then a launch in New York; or data from before a price change.
Also: retrieval
Fetching relevant documents and giving them to an LLM with each question, so it answers from your sources.
ExampleAn HR assistant that answers from the company's policy handbook. This course hub's chat works the same way.
Hiding part of the data and learning to fill it in, so the data labels itself. How LLMs are pre-trained.
Example'The patient missed the ___ because of rain' teaches the model 'appointment' with no human labelling.
Learning from a few labelled examples plus many unlabelled ones, usually because labelling is expensive.
ExampleA million support tickets and the budget to label two thousand.
A single pandas column: values of one type, with a name and a label for each row.
Exampledf['Contract'] is a Series; df[['tenure', 'Churn']] is a DataFrame.
Using data to make one or a few important decisions carefully, with uncertainty on the table.
ExampleDeciding whether to open a new branch, or whether a model is safe to launch.
Also: stratify
Splitting so that each part keeps the same share of each label as the whole dataset.
ExampleAbout 26.5% of customers left; stratify=y keeps that share in both the training and test parts.
Also: not applicable
A category value that means the question doesn't apply, rather than 'no'.
Example'No internet service' in TechSupport: the customer couldn't have tech support at all.
Also: learning with labels
Learning from examples that come with the correct answer attached. The most common and most reliable framing when labels can be had.
ExamplePast customers with a known outcome (left or stayed) teach the model to score current customers.
Also: leakage, data leakage
When a feature contains information that only exists after the outcome. The model looks brilliant in testing and fails in real use.
ExampleLastContactReason: for most leavers the last contact was the cancellation call itself; 96% of 'Cancellation request' customers left.
Also: holdout set, locked test set
Data set aside before exploring or modelling and opened once, at the very end. Anything you look at stops being a fair exam, including for you.
Exampletest_locked.csv is saved in Session 1 and opened in Session 3.
A deliberately plain description of machine learning: it makes many small, repeated decisions by putting a label on each thing that comes in.
ExampleEmail in, 'spam' out. Frame of a game in, 'move left' out. Patient on day 61 in, '34 mg' out.
Also: cut-off, decision threshold
The score above which a prediction counts as positive. Where to put it is a business decision about the cost of each mistake.
ExampleLower the threshold and the team catches more leavers but also calls more customers who would have stayed.
Also: rules-based system
A person writes the recipe (the rules) by hand; the computer applies it to data to get answers.
ExampleA VAT calculation: the rule is known and fixed, so there is nothing to learn.
Also: holdout, train_test_split
Dividing the data into a part the model learns from and a part kept aside to measure how it does on rows it has never seen.
Example80% of customers for training, 20% for testing, stratified and with a fixed seed.
Also: wrong question
Giving the right answer to the wrong question. The most expensive mistake in data science; no algorithm can fix it.
ExampleThe car park model answered 'is this space full?' with 96% accuracy when the real question was 'where are the empty spaces?'.
Also: to_numeric, astype
Turning a column into another type. With errors='coerce', values that can't be converted become NaN instead of stopping the code.
Examplepd.to_numeric(df['TotalCharges'], errors='coerce') turns 11 blank spaces into NaN.
Also: grain
What one row represents. Choosing it is a design decision that changes the dataset and the problem.
ExampleOne row per customer (our snapshot) versus one row per customer per Monday (what weekly scoring really needs).
Learning from data without answers: the algorithm finds structure, such as groups, and people decide what the groups mean.
ExampleClustering photos of two cats may group them by sunny versus cloudy, not by cat.
Also: practice exam
Data used to compare options and tune a model: the practice exams. In this course it is carved out of the training part with cross-validation (Session 2).
ExampleComparing tree depths 3, 4 and 6 on validation data, never on the test set.
Also: vectorized operation
Applying an operation to a whole array or column at once instead of looping over values one by one. Shorter code, far faster.
Examplechurned.mean() instead of a for-loop: about 150 times faster on a million customers.