0%
BuildHow a chatbot works, by building oneabout 28 min, 8 steps

Files, rows, features, and labels

Read a data file, inspect its columns, and load it into a dataset, ready for the table model later in this level.

A file is saved text

This lesson shows how a program reads a table of data from a file. The last module of this level trains a model on exactly this kind of table, and the toy language model in the next lesson reads its text from a file too.

The Files tab in this article's practice editor already contains the teaching files. Select a file to read or edit its text, then go back to Program to run the blocks that read it. Changing a file's text changes the data, not the blocks.

The file practice-deliveries.csv holds four delivery records:

CSVHover over a line to see an explanation
km,drivers,rain,label2,4,0,on-time9,1,1,late4,3,0,on-time8,2,1,late

CSV stands for comma-separated values: plain text in which commas separate the columns. The first line is a header that names four columns. Each later line is one example. The first three values in a row are features, the inputs a model can use. The last value is the label, the answer it should learn. Rain is written as 0 for no and 1 for yes. Write a meaning like that down; a bare 0 does not explain itself.

Build with me · 1

Read a file as text first

Reading the raw file first shows you exactly what the loader will receive: a header line, then one line per delivery, with commas between the values.

The filename ends in .csv so that people recognise the format. That does not make it a program. This block only reads the text; a later load block turns its lines into examples.

Blocks at this stageWorked example
say
the text in filepractice-deliveries.csv

Choose Files beside the article and inspect practice-deliveries.csv. Return to Program and put the file-text value inside say.

What to look for

The header and four data lines print exactly as text.

Make it yours

Change one invented distance in Files and rerun. Observe that the program reads the updated contents rather than an old screenshot.

Read the columns before using them

The first example means a distance of 2 kilometres, four drivers, no rain, and an on-time delivery. It does not mean two miles, or zero millimetres of rain. The units and meanings are part of the data, even though the file only shows numbers.

A prediction input must give the same three features in the same order. 2, 4, 0 matches. 4, 2, 0 swaps distance and drivers and asks a different question, and a model usually cannot tell that you swapped them.

Load the file into a dataset

A file is not yet something a model can learn from. The program has to create a dataset and fill it from the file.

Build with me · 2

Create the destination dataset

The name deliveries belongs to a dataset that this run creates and keeps only while it runs. It is different from the filename practice-deliveries.csv. A file can exist without any dataset having been created yet.

The loader needs a destination so it knows where to add the rows it reads. Creating that destination first makes the order explicit.

Blocks at this stageWorked example
make a dataset calleddeliveries
sayBefore loadthen
how many examples indeliveries

Add make a dataset called deliveries, then print its size. Keep the filename out of the dataset-name field.

What to look for

Before load is 0 even though the file already contains four rows.

Make it yours

Explain why the presence of a CSV in Files does not automatically train a model.

One thing to know before the next stage. The load block asks which column holds the label, and it counts columns from zero: column 0 is the first column, column 1 the second, and so on. Negative numbers count back from the end, so column -1 is the last column. In km,drivers,rain,label, column 0 is km and column -1 is label. The list block you used earlier counts its items from 1 instead, so whenever a block asks for a position, check which way it counts.

Build with me · 3

Choose the answer column and header

The label column -1 means the last column, counting from the end as explained above. Here that is label, whose values are on-time and late. Skip the first row because its words are column names, not a delivery. The other three columns become inputs.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadpractice-deliveries.csvintodeliveriesusing column-1as the label, andskip the first row
sayLoadedthen
how many examples indeliveries

Connect load practice-deliveries.csv into deliveries. Enter -1 for label column and choose skip the first row.

What to look for

Loaded is 4. The loader also reports four imported rows.

Make it yours

Inspect which field would become the answer if you used column 0. Explain why that would define a different prediction task.

Build with me · 4

Check meaning as well as count

The first row supplies features [2, 4, 0] and label on-time. Interpret it aloud: two kilometres, four drivers, no rain, outcome on-time. Units and the 0/1 code for rain are part of the data's meaning.

The values are numeric inputs after loading, while the outcome remains a label. Correct row count alone would not reveal a swapped column or wrong interpretation of rain.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadpractice-deliveries.csvintodeliveriesusing column-1as the label, andskip the first row
show datasetdeliveries

Add show dataset deliveries below load. Compare the printed row with its source line in Files.

What to look for

Four feature/label pairs print in source order.

Make it yours

Swap the first two numbers in one CSV row and inspect the change. Restore it after explaining how the row's meaning changed.

Build with me · 5

Build a prediction-shaped list

A list packages the three intended inputs in the same order as training: kilometres, drivers, rain. It contains three values, not one string saying 2, 4, 0. The list value block expresses that structure directly.

An ID identifies a row for later matching. It is not a fourth numerical feature in this task. Including an ID in a numeric distance calculation would change or break what the model is comparing.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadpractice-deliveries.csvintodeliveriesusing column-1as the label, andskip the first row
setfeaturesto
a list of2, 4, 0
sayPrediction inputthen
features

Put a list-of block in set features. Enter 2, 4, 0 as its items and print the resulting variable.

What to look for

Prediction input shows a three-number list.

Make it yours

Choose another fictional delivery and state each feature's meaning before editing its number.

Files with no answers

A file you are asked to predict for usually has no label column, because the labels are what you have to supply. It often has an ID column instead. An ID such as practice01 names a row, so that each prediction can be matched back to it later, just like the IDs in your Level 2 predictions file. It is a name, not a measurement, so it must never be used as a feature.

Build with me · 6

Load targets whose answers are unknown

The new file includes IDs and features, but no correct answer column. Enter none so the loader does not silently take rain as the label. It keeps all source fields available for inspection and output matching.

Because answers are absent, predictions can be produced later, but accuracy cannot be calculated honestly for these rows. An empty placeholder label is not a correct answer.

Blocks at this stageWorked example
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
sayTarget rowsthen
how many examples inrows
show datasetrows

Create rows and load couriers-new.csv with label column none and header skipped. Inspect the result.

What to look for

Target rows is 3; each row retains its ID and numerical values.

Make it yours

Compare the header with the labelled file's header. Identify exactly which column is missing and which column was added.

Build with me · 7

Select only the intended features

The simple block loader retains every non-label column. In this teaching target file that includes the ID, so we build the three numeric feature lists explicitly. Their order must match the source rows exactly.

This manual approach is understandable for three rows and unsuitable for hundreds. Level 4 replaces it with named-column selection in Python. Here it makes the alignment requirement visible without pretending an ID is a measurement.

Blocks at this stageWorked example
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
saySource rowsthen
how many examples inrows
sayNumeric inputsthen
how many things ininputs
for eachfeaturesin listinputs
say
features

Build inputs from the three rows, leaving out each ID. Compare each list against the matching source line in Files before running.

What to look for

Both counts are 3, and the three lists match the numerical portions of the target rows in order.

Make it yours

Deliberately swap two lists, explain which IDs would receive the wrong predictions, then Undo before continuing.

When loading goes wrong

practice-deliveries.csv is the name of a file. deliveries is the name of a dataset that your program creates while it runs. The load block connects them. If you change the file, run the program again to rebuild the dataset from the new text.

If loading says it cannot find the file, check the spelling, capital letters, and the .csv ending, and check that the file is in this editor's Files tab. This teaching loader splits every line at every comma, so keep commas out of the values themselves.

Build with me · 8

Inspect both collections before modelling

This final check brings together file contents, dataset counts, feature meaning, and target alignment. Do it before choosing a model. A model cannot recover the intended meaning of a column you accidentally swapped.

If a filename is wrong, fix the name or file. If a count is wrong, inspect header handling and source rows. If the values are wrong, inspect feature order and units. These are different errors and deserve different repairs.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadpractice-deliveries.csvintodeliveriesusing column-1as the label, andskip the first row
sayLabelled examplesthen
how many examples indeliveries
show datasetdeliveries
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
sayTarget rowsthen
how many examples inrows
sayNumeric inputsthen
how many things ininputs

Run the combined inspection. Use Files to compare any unexpected row, then rerun after a deliberate correction.

What to look for

The supplied files yield four labelled examples, three target rows, and three numeric target inputs.

Make it yours

Write one sentence describing this data: the inputs in order, their units, what 0 and 1 mean for rain, the answer labels, and the separate job of IDs.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calleddeliveries
loadpractice-deliveries.csvintodeliveriesusing column-1as the label, andskip the first row
sayLabelled examplesthen
how many examples indeliveries
show datasetdeliveries
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
sayTarget rowsthen
how many examples inrows
sayNumeric inputsthen
how many things ininputs
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  deliveries = new_dataset("deliveries")load_csv(deliveries, "practice-deliveries.csv", -1, True)print("Labelled examples", deliveries.size())deliveries.show()rows = new_dataset("rows")load_csv(rows, "couriers-new.csv", "none", True)inputs = []inputs.append([3, 4, 0])inputs.append([9, 1, 1])inputs.append([5, 2, 1])print("Target rows", rows.size())print("Numeric inputs", len(inputs))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in