0%
BuildRound 3 rehearsal: predict from a tableabout 37 min, 10 steps

Train a model on a table

Load a table, hold rows back, train a closest-example model and a baseline, and export one prediction per row ID.

The file you will train on

This module ends with the kind of file a competition round scores: one predicted label for each row ID. This lesson builds the whole pipeline on a small table, from a data file to that final file, with every step in the editor beside the article.

The Files tab includes couriers-practice.csv. Distance is in kilometres, drivers is how many drivers were working, and rain is 0 for no and 1 for yes.

CSVHover over a line to see an explanation
km,drivers,rain,label1,4,0,on-time2,4,0,on-time3,3,0,on-time4,3,0,on-time5,4,0,on-time6,4,0,on-time2,2,1,on-time7,1,1,late8,1,1,late9,2,1,late10,1,0,late8,2,0,late

These twelve invented deliveries are for learning the steps, not for real delivery forecasting. Seven are on time and five are late. There is no ID column, so all three columns before the label are features.

Build with me · 1

Create and load the labelled dataset

Create deliveries before loading into it, just as in the files lesson. Choosing the last column as the label leaves exactly the three intended inputs.

Check the count and the preview before building any model. A wrong count is much easier to fix now than after training.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
sayLoadedthen
how many examples indeliveries
show datasetdeliveries

Open couriers-practice.csv in the Files tab. Build make dataset, load with -1 and skipped header, then labelled size and show.

What to look for

Loaded is 12; the file contains seven on-time and five late labels.

Make it yours

Find one row of each class and explain its input values. Do not treat those invented associations as a real causal rule.

Hold some rows back

As in Level 2, a model has to be tested on rows it did not learn from, so the next stage sets some aside before any model exists.

Build with me · 2

Split before choosing a model

The split keeps 75 percent for training and the remainder for development. Seventy-five percent of twelve is nine, leaving three. The block shuffles the rows in the same fixed way every time, so another Run with the same file produces the same split.

Training rows fit the model. Development rows help judge and choose versions. Once you use those results to make decisions, they are not an untouched final test.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
sayTraining rowsthen
how many examples intraining
sayDevelopment rowsthen
how many examples indevelopment

Put split deliveries into training and development after loading. Set 75 in its percentage socket and print both sizes.

What to look for

Training rows is 9 and Development rows is 3.

Make it yours

Explain why repeatedly pressing Run here does not test many different random splits.

How a closest-example model decides

The model you are about to create is called closest example. Its idea is simple: to predict a new delivery, find the training deliveries most like it and let them vote. Its usual name is k nearest neighbours, often shortened to KNN, where k is how many neighbours vote. Here k is 3.

"Most like it" needs a number, so the model measures a distance between two rows:

  1. Subtract one row's values from the other's, column by column.
  2. Square each difference, which means multiplying it by itself. This makes every difference positive, so a difference of 2 and a difference of -2 count the same.
  3. Add the squares, then take the square root of the total. That is the distance: one number saying how far apart the two deliveries are.

This is the ordinary straight-line distance you would measure with a ruler between two points on a map, with three columns instead of two.

Take the new delivery you will predict later in this lesson: 6 km, 2 drivers, rain 1. The model compares it with every row in the training set, and never with the development rows it has not been given. Here are five of the nine training rows, with the size of each difference:

Training row (km, drivers, rain)LabelDifferencesSquares addedDistance
7, 1, 1late1, 1, 01 + 1 + 0 = 2square root of 2, about 1.41
8, 1, 1late2, 1, 04 + 1 + 0 = 5square root of 5, about 2.24
6, 4, 0on-time0, 2, 10 + 4 + 1 = 5square root of 5, about 2.24
4, 3, 0on-time2, 1, 14 + 1 + 1 = 6square root of 6, about 2.45
9, 2, 1late3, 0, 09 + 0 + 0 = 9square root of 9, exactly 3

The other four training rows are all at least 2.45 away. So the three closest rows are the first three: late, late, and on-time. Count the votes: late 2, on-time 1. The model predicts late, and its score is the winning share of the vote, 2 out of 3, which it shows as 66.7. You will see exactly this when you run the prediction stage.

Now notice what the distance ignores: it treats every column the same, whatever the column means. That only works when the columns use numbers of similar size. Suppose distance had been recorded in metres instead. The new delivery would be 6,000 and the closest row 7,000, a difference of 1,000, which squares to 1,000,000. A difference of 2 drivers adds only 4, and rain at most 1. Distance would swamp the other two columns, and the model would in effect ignore drivers and rain. Even in kilometres, which run from 1 to 10, distance can count for more than rain, which is only ever 0 or 1. Level 4 shows how to rescale columns to similar ranges before measuring distance.

Now build it.

Build with me · 3

Choose the closest-example rule

The closest-example model predicts by finding the three training rows nearest to a new input and letting them vote, exactly as in the worked example above. It measures distance on the raw numbers, so a column with much larger numbers counts for more.

Creating the model does not look at any data yet. Training, in the next stage, gives it the rows it will search.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
show whatcourierlearned

Choose closest example in make-model and name it courier. Inspect it before fitting.

What to look for

The model report shows the nearest model type and zero training rows.

Make it yours

Redo the first row of the worked example with distance in metres: 6,000 and 7,000. Which column now decides the distance?

Build with me · 4

Train only on the training rows

Training links courier to the nine training examples. It must not use development answers. A later score on development then asks how the model handles rows it was not trained on.

For a closest-example model, training simply means keeping the training rows so it can search them later. Other models, like your image models, learn numbers instead. Either way, training means preparing the model from data.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
show whatcourierlearned

Connect train courier on training. Check that the dataset menu says training, not deliveries or development.

What to look for

The model reports nine training examples.

Make it yours

Point to the exact block that would leak development information if you selected the full dataset there.

Measure it against a baseline

A score means little on its own. The comparison is the baseline you met in Level 2: always predict the most common label in the training rows.

Build with me · 5

Measure on held-out labelled examples

Accuracy compares predicted labels with known correct labels. Each correct example contributes one to a count, divided by three and multiplied by one hundred. One changed outcome therefore moves this tiny score by about 33.3 percentage points.

A high score on three examples is weak evidence about future deliveries. Its first job here is to check that training on some rows and testing on others works.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
sayModel accuracythen
out of 100, the accuracy ofcourierondevelopment

Put the accuracy value for courier on development inside a labelled say. Keep it below fitting.

What to look for

Model accuracy is 100.0: all three development rows are predicted correctly. The score comes from the actual model, not a preset answer.

Make it yours

Describe the result using a count out of three as well as a percentage. State the small-sample limitation.

Build with me · 6

Prepare a simple reference fairly

The always-most-common-label model learns its one answer from the same training rows. It ignores feature values and predicts that constant for every input. This simple reference can perform surprisingly well when labels are uneven.

Both models are evaluated on the same development rows. Choosing the constant from development labels would peek at the answers used to judge it, so the train-baseline block must also say training.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayTraining majoritythen
the most common label intraining
sayModel accuracythen
out of 100, the accuracy ofcourierondevelopment
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment

Create baseline with always the most common label, train it on training, and print both accuracies on development.

What to look for

Training majority is on-time. Model accuracy is 100.0 and Baseline accuracy is 66.7: always saying on-time gets two of the three development rows right.

Make it yours

Explain what an accuracy tie would mean here, and why it would not show the two models made identical predictions on all possible inputs.

With only three development rows, these scores show that the pipeline works, not that the model is good. The next lesson looks at baselines properly.

Now ask the model about the delivery from the worked example, and check its answer against your hand calculation.

Build with me · 7

Predict one correctly shaped input

The new list means six kilometres, two drivers, rain. A three-number list matches the training rows' feature order. The score is the winning share of the three votes, so with two labels it can only be 66.7 or 100. It is not the chance that the delivery will really be late.

This row has no supplied correct answer, so its prediction is not an accuracy measurement. Keep the two kinds of output separate in your explanation.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayModel accuracythen
out of 100, the accuracy ofcourierondevelopment
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment
setfeaturesto
a list of6, 2, 1
sayFeaturesthen
features
sayPredictionthen
whatcouriersays about
features
sayVote sharethen
out of 100, how surecourieris about
features

Put a list-of value in set features. Use the same features variable for prediction and confidence, then print both.

What to look for

After the development scores, the program prints the three feature values, the prediction late, and a vote share of 66.7, matching the hand calculation above.

Make it yours

Choose another fictional delivery inside the observed distance range. Predict which nearby examples might influence it, then compare.

Predict a file of unlabelled rows

The Files tab also includes couriers-new.csv, three deliveries whose outcomes are unknown:

CSVHover over a line to see an explanation
id,km,drivers,rainpractice01,3,4,0practice02,9,1,1practice03,5,2,1

Each row has an ID, which names the row, and three features. The ID has to stay attached to its row's prediction, but it must never be used as a feature: practice02 is not a distance.

Build with me · 8

Keep target IDs out of model inputs

The target CSV carries an ID plus three features, but no labels. Loading it with none preserves those fields for output matching. The three explicit numerical lists leave out IDs and follow the source order.

Equal counts are needed but not enough: three wrong or reordered inputs still produce three outputs. Compare each list with its matching file row before predicting.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
sayTarget countthen
how many examples inrows
sayInput countthen
how many things ininputs

Load couriers-new.csv into rows with label column none. Build inputs in the exact file order, using only km, drivers, rain.

What to look for

Target count and Input count are both 3.

Make it yours

Explain why practice02 belongs in the exported output but not inside the numerical distance calculation.

Build with me · 9

Append one prediction on each loop turn

The loop item features is one three-number input at a time. Predict courier on that item, then append the returned label to predictions. Appending preserves loop order.

Create the empty predictions list before the loop. Putting the make-list block inside the loop would erase earlier predictions and leave only the last label. This is the same running-state principle as the earlier total example.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
make an empty list calledpredictions
for eachfeaturesin listinputs
add
whatcouriersays about
features
topredictions
sayPrediction countthen
how many things inpredictions
show listpredictions

Create predictions, then for each features in inputs, place add prediction-to-predictions inside its body. Print the final count below the loop.

What to look for

Prediction count is 3 and the output list contains three labels in target order.

Make it yours

Trace the first two turns by hand, naming the current input and the list after each append.

Build with me · 10

Match IDs to predictions in the final CSV

The submission operation pairs each prediction with the corresponding source row's ID. The header is id,label, followed by exactly three data rows. It cannot repair a feature list you built in the wrong order, so alignment checks come first.

The development scores describe known labelled examples. These target rows are unlabelled, so do not call their predictions correct merely because a CSV was produced. A valid format and a useful model are separate achievements.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayModel accuracythen
out of 100, the accuracy ofcourierondevelopment
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
make an empty list calledpredictions
for eachfeaturesin listinputs
add
whatcouriersays about
features
topredictions
sayTarget countthen
how many examples inrows
sayPrediction countthen
how many things inpredictions
savepredictionsas a submission forrows

Add save predictions as a submission for rows after the loop. Inspect the printed CSV section, then use Download submission.csv to keep the generated result. The downloaded file includes only the header and result rows, not the surrounding status messages.

What to look for

The CSV has practice01, practice02, and practice03 exactly once in original order, each paired with a predicted label.

Make it yours

Explain the complete pipeline from source file through feature order, split, fitting, comparison, prediction, and row matching. Use the collapsed reference to check a missing connection.

Explain what you can claim

The new rows have no known labels, so you cannot report their accuracy. Your development score compares the two models on three known examples. It shows the pipeline works; it is not evidence of dependable delivery predictions. Keep those two statements separate when you write up a result.

If the counts are wrong, check the loading and the loops before changing the model. If the format is right but the predictions are weak, look at the data, the features, the scale of each column, and the choice of model. Keeping those two kinds of problem apart will help in the rehearsal, which brings its own data and rules.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotraininganddevelopmentkeeping75percent for training
make aclosest examplemodel calledcourier
traincourierontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayModel accuracythen
out of 100, the accuracy ofcourierondevelopment
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment
make a dataset calledrows
loadcouriers-new.csvintorowsusing columnnoneas the label, andskip the first row
make an empty list calledinputs
add
a list of3, 4, 0
toinputs
add
a list of9, 1, 1
toinputs
add
a list of5, 2, 1
toinputs
make an empty list calledpredictions
for eachfeaturesin listinputs
add
whatcouriersays about
features
topredictions
sayTarget countthen
how many examples inrows
sayPrediction countthen
how many things inpredictions
savepredictionsas a submission forrows
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set(    "a an and are as at be been but by can could did do for from had has have "    "he her his i if in is it its me my not of on or our she so than that the "    "their them then there they this to too us was we were what when which "    "who will with would you your".split())  def _content_words(value):    words = _words_in(value)    kept = [word for word in words if word not in _EVERYDAY_WORDS]    return kept if kept else words  def _as_numbers(value):    if isinstance(value, (list, tuple)):        return [float(item) for item in value]    return [float(value)]  def _is_numeric(value):    try:        _as_numbers(value)        return True    except (TypeError, ValueError):        return False  def _distance(left, right):    """How far apart two examples are. Numbers use straight line distance,    text uses how many words the two do not share."""    if _is_numeric(left) and _is_numeric(right):        a = _as_numbers(left)        b = _as_numbers(right)        while len(a) < len(b):            a.append(0.0)        while len(b) < len(a):            b.append(0.0)        total = 0.0        for i in range(len(a)):            total += (a[i] - b[i]) ** 2        return math.sqrt(total)    a = set(_words_in(left))    b = set(_words_in(right))    if not a and not b:        return 0.0    shared = len(a & b)    return 1.0 - (shared / float(len(a | b)))  class Model:    """Three small classifiers behind one name.     nearest  looks for the closest example it was trained on    words    scores how many words the input shares with each label    common   always answers with the most common label, the baseline to beat    """     def __init__(self, kind="nearest", name="model"):        self.kind = kind        self.name = name        self.data = None        self.word_scores = {}     def train(self, dataset):        self.data = dataset        self.word_scores = {}        if self.kind == "words":            for row in dataset.rows:                bucket = self.word_scores.setdefault(row["label"], {})                for word in _content_words(row["features"]):                    bucket[word] = bucket.get(word, 0) + 1        print("Trained " + self.name + " on " + str(dataset.size()) + " examples.")     def _scores(self, features):        if self.data is None or not self.data.rows:            return {}        if self.kind == "common":            counts = {}            for row in self.data.rows:                counts[row["label"]] = counts.get(row["label"], 0) + 1            return counts        if self.kind == "words":            # Score each label by how much of its training vocabulary shows up            # in the input, divided by how much text that label was trained on            # so a label with more examples cannot win on volume alone.            labels = self.data.labels()            scores = {}            for label in labels:                bucket = self.word_scores.get(label, {})                seen = sum(bucket.values()) or 1                running = 0.0                for word in _content_words(features):                    running += bucket.get(word, 0) / float(seen)                scores[label] = round(running, 6)            if sum(scores.values()) == 0:                # Nothing in the input was ever seen in training. Say so by                # splitting the vote evenly, which reads as low confidence.                return dict((label, 1) for label in labels)            return scores        ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"]))        neighbours = ranked[: min(3, len(ranked))]        scores = {}        for row in neighbours:            scores[row["label"]] = scores.get(row["label"], 0) + 1        return scores     def predict(self, features):        scores = self._scores(features)        if not scores:            return ""        return max(scores, key=lambda label: scores[label])     def confidence(self, features):        """How much of the vote the winning label took, out of 100."""        scores = self._scores(features)        total = sum(scores.values())        if not scores or total == 0:            return 0.0        best = max(scores.values())        return round(100.0 * best / float(total), 1)     def accuracy(self, dataset):        if not dataset.rows:            return 0.0        right = 0        for row in dataset.rows:            if self.predict(row["features"]) == row["label"]:                right += 1        return round(100.0 * right / float(len(dataset.rows)), 1)     def show(self):        print(self.name + " is a " + self.kind + " model.")        if self.data is None:            print("  It has not been trained yet.")            return        print("  Trained on " + str(self.data.size()) + " examples.")        print("  Labels it can answer with: " + ", ".join(self.data.labels()))  def new_model(kind="nearest", name="model"):    return Model(kind, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  deliveries = new_dataset("deliveries")load_csv(deliveries, "couriers-practice.csv", -1, True)training, development = deliveries.split(75)courier = new_model("nearest", "courier")courier.train(training)baseline = new_model("common", "baseline")baseline.train(training)print("Model accuracy", courier.accuracy(development))print("Baseline accuracy", baseline.accuracy(development))rows = new_dataset("rows")load_csv(rows, "couriers-new.csv", "none", True)inputs = []inputs.append([3, 4, 0])inputs.append([9, 1, 1])inputs.append([5, 2, 1])predictions = []for features in inputs:    predictions.append(courier.predict(features))print("Target count", rows.size())print("Prediction count", len(predictions))save_submission(predictions, rows)

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in