0%
BuildYour first AI appabout 36 min, 10 steps

Is it a question? Train a text model and gate it

Create a dataset, create a text model, train it, predict a label, and add a branch for when it is unsure.

The program you will build

The program reads a message, predicts whether it is a question or a statement, and then either replies or says it is unsure. You will build it in ten small stages beside this article. Each stage is a complete program: Try this stage loads everything it needs, and training happens in your browser when you press Run.

The steps are the ones you followed in Level 2: collect labelled examples, create a model, train it, ask it about a new input, then decide what to do with its answer. First we get the program working. Then we find where it fails. This model matches words; it does not understand English grammar.

Collect labelled examples

A model learns from examples, so the program starts with somewhere to keep them.

Build with me · 1

Create the collection before filling it

The dataset named examples begins empty. It will hold pairs: some input text and the label we want that text to have. Creating a dataset is bookkeeping; it does not create a model or teach one anything.

Decide the meaning of your labels before collecting rows. For this teaching task, question includes a request for an answer, even without a question mark. Statement means it states something. The model will count word associations, not apply this definition as a grammar rule.

Blocks at this stageWorked example
make a dataset calledexamples
sayExamplesthen
how many examples inexamples

From Data, add make a dataset called examples. Put its size value inside a labelled output below it.

What to look for

Examples is 0.

Make it yours

Rename the dataset in its make-dataset block and in every menu that uses it. The name can change without changing what a question means.

Build with me · 2

Add one input with its answer

An example has two parts. Please explain gravity is the input. question is the target answer supplied by a person. The label is not another input word that the model is allowed to read during prediction.

One example is enough to inspect the mechanics, but not enough to build a useful recogniser. At this stage we only want proof that the add block ran on the intended dataset.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
sayExamplesthen
how many examples inexamples
show datasetexamples

Connect add to examples below the make-dataset block. Type the phrase in features and question in label. Connect the size and show blocks underneath.

What to look for

Examples is 1, and the shown row pairs the phrase with question.

Make it yours

Change the phrase to another request, keeping the label spelling exactly question. Explain why Question would be a different label.

Build with me · 3

Add a contrasting label

A classifier must have alternatives to compare. A second label gives it two possible answers, but the topics are also different: gravity appears only in a request and parcel only in a statement. That accidental connection will matter later.

Notice that both examples are added to the same dataset. Creating a new dataset before every add would erase the previous collection, leaving only the last row when training begins.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
sayExamplesthen
how many examples inexamples
show datasetexamples

Add the second example underneath the first, selecting the existing examples dataset. Keep exactly one make-dataset statement at the top.

What to look for

Examples is 2 and both labels appear in the printed rows.

Make it yours

Write a question about a parcel and a statement about gravity on paper. Keep them for later probes rather than adding them immediately.

Here is the full teaching collection. It follows the rule from the first stage, so a request such as "Please explain gravity" counts as a question even without a question mark. Spell each label exactly the same way every time.

Example textLabel
Please explain gravityquestion
Could you explain magnetsquestion
Why does ice meltquestion
How does a battery workquestion
Where can I find the libraryquestion
Can you describe a volcanoquestion
What causes thunderquestion
How do seeds growquestion
The parcel arrived safelystatement
My parcel arrived yesterdaystatement
Dinner tastes deliciousstatement
Our team won todaystatement
The blue door is closedstatement
Music played softlystatement
The dog slept peacefullystatement
We painted the fencestatement

Build with me · 4

Grow the collection deliberately

Here are sixteen labelled teaching sentences. Eight request answers and eight state something. Repetition across words such as explain or parcel makes an association the toy model can count. Variety helps reveal how it uses words, but this is still a tiny invented dataset.

The size checkpoint comes after all add statements. The show operation may display only a preview, so count and preview answer different questions: how many rows exist, and what some rows actually contain.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
sayExamplesthen
how many examples inexamples
show datasetexamples

Build or load the sixteen-row stack. Read the label on each add block. Run and check the final count before inserting any model block.

What to look for

Examples is 16. The preview shows input/label pairs, not a learned model.

Make it yours

Replace one sentence with your own under the same definition. Keep the class counts balanced during this first controlled comparison.

Create a model, then train it

The examples are ready. Now the program needs a model to learn from them.

Build with me · 5

Choose a model and give it a name

The make-model block creates judge and selects word matching as its algorithm. An algorithm is the procedure the model uses to learn and predict. This one compares word associations. Choosing it does not automatically connect it to examples.

The dataset and model have different names because they have different jobs. examples stores labelled rows; judge will store the learned information. You can inspect judge before training to see that creating a model is not training it.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
show whatjudgelearned

From Model, choose make a word matching model called judge. Put it below the examples. Add show what judge learned after it.

What to look for

The model report identifies a words model with zero training examples. No prediction is made yet.

Make it yours

Locate every occurrence of examples and judge in the stack. Explain what each menu is asking you to choose.

Build with me · 6

Connect the model to the collection

Train judge on examples is the instruction that reads the labelled collection and records its word patterns in the model. It must run after collection and after model creation.

Putting another make judge block below training would replace the trained model with an empty one. The same thing happens in Python later: creating a new model under the same name throws away the trained one.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
trainjudgeonexamples
show whatjudgelearned

Insert train judge on examples between the make-model block and its show block. Keep the model and dataset menus in the correct positions.

What to look for

The training message reports 16 examples, and the model report now says it was trained on 16 examples.

Make it yours

Temporarily move a fresh make-model block after training and inspect the reset model, then Undo. Do not diagnose its missing learning by adding more examples.

The order matters: make the dataset, add the examples, make the model, train the model. A second make-dataset block after the examples would empty the collection again, in the same way that a second make-model block after training empties the model.

Use the trained model

A trained model can now be given a message it has never seen. Asking it for a label is prediction, which Level 2 also called inference. The model does not learn anything from the new message; it only answers.

Build with me · 7

Give the trained model a new message

Prediction uses the trained model. It does not add the new message to training, and it does not require that you provide the answer. Store the input under message so the printed input and predicted label refer to the same text.

This familiar wording is a wiring check: if the program reads the wrong value, fix that first. A successful familiar case is not evidence that the model understands new English sentences.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
trainjudgeonexamples
setmessagetoPlease explain magnets
sayMessagethen
message
sayPredictionthen
whatjudgesays about
message

Nest the rounded message variable in what judge says about. Put that prediction value inside a labelled say, after training.

What to look for

The supplied message is printed and the model returns question.

Make it yours

Replace the message with The parcel arrived safely and predict the output. Then try new wording on the same topic.

Build with me · 8

Keep the label and its score together

The model supplies two things: a winning label, and a score out of 100 saying how much of its evidence pointed to that label. We store each because the next decision will need them. Both blocks read the same message and the same model.

These unfamiliar words have no useful overlap with the training vocabulary. This implementation shares support evenly when no words match, giving 50 for two classes. A tie still returns a label; the application must decide whether to present it as a useful answer.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
trainjudgeonexamples
setmessagetomarble lantern violin
setlabelto
whatjudgesays about
message
setsureto
out of 100, how surejudgeis about
message
sayPredictionthen
label
sayScorethen
sure

Add assignments for label and sure. Place prediction and confidence value blocks in their sockets, both reading message. Print the stored values underneath.

What to look for

Score is 50 for this no-overlap input. Treat the tie's label as an arbitrary result of the implementation, not a recognition success.

Make it yours

Change message to a known phrase. Compare the score, then explain why a larger number alone cannot prove correctness.

A confidence gate

The model always returns a label, even for input unlike anything it trained on, because question and statement are the only answers it has. The program around the model decides whether that label deserves to be shown. A rule that checks the score before showing the label is called a confidence gate.

Build with me · 9

Decide how the application handles uncertainty

The gate is an application rule you write. If sure is below 60, it prints an unsure message. Otherwise, a second if chooses the question or statement response. The verdict messages must live inside the appropriate branches.

The threshold is a trial choice, not a proven safe boundary. A score of exactly 60 passes because the condition is less than 60. Raising the threshold means more unsure answers and fewer accepted ones; it cannot catch every confidently wrong answer.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
trainjudgeonexamples
setmessageto
the line typed in
setlabelto
whatjudgesays about
message
setsureto
out of 100, how surejudgeis about
message
if
sure
<60
then
sayI am not sure. Please rephrase.
otherwise
if
label
isquestion
then
sayThat looks like a question.
otherwise
sayThat looks like a statement.
sayRecorded scorethen
sure

Build the outer gate and inner label check. Press Run, answer the waiting Terminal with an unfamiliar message, and keep the score output outside the gate for diagnosis.

What to look for

After answering marble lantern violin, the program prints the unsure response and a score of 50.

Make it yours

Try thresholds 50, 60, and 80 with the same Terminal answer. For each, count the wrong answers it accepted and the times it said unsure.

The gate catches messages the model has no evidence about. It cannot catch a message the model is sure about and gets wrong.

Build with me · 10

Expose a confident mistake

Gravity is interesting is a statement. Yet gravity appeared in a question example, and this matcher ignores the sentence structure that would help distinguish the two. It can confidently choose question.

Keep the raw label and score visible during development so a reassuring response cannot hide how the decision arose. A program can run every block correctly and still make a poor prediction, because the model only counts words and learned from sixteen examples.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
trainjudgeonexamples
setmessageto
the line typed in
setlabelto
whatjudgesays about
message
setsureto
out of 100, how surejudgeis about
message
sayInputthen
message
if
sure
<60
then
sayI am not sure. Please rephrase.
otherwise
if
label
isquestion
then
sayThat looks like a question.
otherwise
sayThat looks like a statement.
sayRaw labelthen
label
sayRaw scorethen
sure

Press Run and enter Gravity is interesting when Terminal waits for an answer. Do not change the training collection or gate. Compare the intended label with the actual result before making improvements.

What to look for

The example demonstrates a confident question prediction for a statement. The 60-point gate does not prevent this error.

Make it yours

Try your saved parcel question too. Add paired topics only after recording this first version, retrain, and check whether the change helps the held-out wording.

This is a limit of the model and its examples, not a broken if block. The next lesson opens the model up and shows exactly why it happens.

Test it fairly

The messages you just tried were picked to show how the program works, after you had seen the training examples. They are not a fair test. To measure the model, write a separate development set, as you did in Level 2: new wording, with the same topic written both as a question and as a statement. For each message, record the true label, the predicted label, the score, and whether the gate said unsure. Count the unsure answers too; leaving them out hides how often the program refuses to answer.

Changing the training examples needs another Run, because every Run starts at the top of the program and rebuilds the dataset and the model.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calledexamples
add toexamplesthe examplePlease explain gravitylabeledquestion
add toexamplesthe exampleCould you explain magnetslabeledquestion
add toexamplesthe exampleWhy does ice meltlabeledquestion
add toexamplesthe exampleHow does a battery worklabeledquestion
add toexamplesthe exampleWhere can I find the librarylabeledquestion
add toexamplesthe exampleCan you describe a volcanolabeledquestion
add toexamplesthe exampleWhat causes thunderlabeledquestion
add toexamplesthe exampleHow do seeds growlabeledquestion
add toexamplesthe exampleThe parcel arrived safelylabeledstatement
add toexamplesthe exampleMy parcel arrived yesterdaylabeledstatement
add toexamplesthe exampleDinner tastes deliciouslabeledstatement
add toexamplesthe exampleOur team won todaylabeledstatement
add toexamplesthe exampleThe blue door is closedlabeledstatement
add toexamplesthe exampleMusic played softlylabeledstatement
add toexamplesthe exampleThe dog slept peacefullylabeledstatement
add toexamplesthe exampleWe painted the fencelabeledstatement
make aword matchingmodel calledjudge
trainjudgeonexamples
setmessageto
the line typed in
setlabelto
whatjudgesays about
message
setsureto
out of 100, how surejudgeis about
message
sayInputthen
message
if
sure
<60
then
sayI am not sure. Please rephrase.
otherwise
if
label
isquestion
then
sayThat looks like a question.
otherwise
sayThat looks like a statement.
sayRaw labelthen
label
sayRaw scorethen
sure
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set(    "a an and are as at be been but by can could did do for from had has have "    "he her his i if in is it its me my not of on or our she so than that the "    "their them then there they this to too us was we were what when which "    "who will with would you your".split())  def _content_words(value):    words = _words_in(value)    kept = [word for word in words if word not in _EVERYDAY_WORDS]    return kept if kept else words  def _as_numbers(value):    if isinstance(value, (list, tuple)):        return [float(item) for item in value]    return [float(value)]  def _is_numeric(value):    try:        _as_numbers(value)        return True    except (TypeError, ValueError):        return False  def _distance(left, right):    """How far apart two examples are. Numbers use straight line distance,    text uses how many words the two do not share."""    if _is_numeric(left) and _is_numeric(right):        a = _as_numbers(left)        b = _as_numbers(right)        while len(a) < len(b):            a.append(0.0)        while len(b) < len(a):            b.append(0.0)        total = 0.0        for i in range(len(a)):            total += (a[i] - b[i]) ** 2        return math.sqrt(total)    a = set(_words_in(left))    b = set(_words_in(right))    if not a and not b:        return 0.0    shared = len(a & b)    return 1.0 - (shared / float(len(a | b)))  class Model:    """Three small classifiers behind one name.     nearest  looks for the closest example it was trained on    words    scores how many words the input shares with each label    common   always answers with the most common label, the baseline to beat    """     def __init__(self, kind="nearest", name="model"):        self.kind = kind        self.name = name        self.data = None        self.word_scores = {}     def train(self, dataset):        self.data = dataset        self.word_scores = {}        if self.kind == "words":            for row in dataset.rows:                bucket = self.word_scores.setdefault(row["label"], {})                for word in _content_words(row["features"]):                    bucket[word] = bucket.get(word, 0) + 1        print("Trained " + self.name + " on " + str(dataset.size()) + " examples.")     def _scores(self, features):        if self.data is None or not self.data.rows:            return {}        if self.kind == "common":            counts = {}            for row in self.data.rows:                counts[row["label"]] = counts.get(row["label"], 0) + 1            return counts        if self.kind == "words":            # Score each label by how much of its training vocabulary shows up            # in the input, divided by how much text that label was trained on            # so a label with more examples cannot win on volume alone.            labels = self.data.labels()            scores = {}            for label in labels:                bucket = self.word_scores.get(label, {})                seen = sum(bucket.values()) or 1                running = 0.0                for word in _content_words(features):                    running += bucket.get(word, 0) / float(seen)                scores[label] = round(running, 6)            if sum(scores.values()) == 0:                # Nothing in the input was ever seen in training. Say so by                # splitting the vote evenly, which reads as low confidence.                return dict((label, 1) for label in labels)            return scores        ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"]))        neighbours = ranked[: min(3, len(ranked))]        scores = {}        for row in neighbours:            scores[row["label"]] = scores.get(row["label"], 0) + 1        return scores     def predict(self, features):        scores = self._scores(features)        if not scores:            return ""        return max(scores, key=lambda label: scores[label])     def confidence(self, features):        """How much of the vote the winning label took, out of 100."""        scores = self._scores(features)        total = sum(scores.values())        if not scores or total == 0:            return 0.0        best = max(scores.values())        return round(100.0 * best / float(total), 1)     def accuracy(self, dataset):        if not dataset.rows:            return 0.0        right = 0        for row in dataset.rows:            if self.predict(row["features"]) == row["label"]:                right += 1        return round(100.0 * right / float(len(dataset.rows)), 1)     def show(self):        print(self.name + " is a " + self.kind + " model.")        if self.data is None:            print("  It has not been trained yet.")            return        print("  Trained on " + str(self.data.size()) + " examples.")        print("  Labels it can answer with: " + ", ".join(self.data.labels()))  def new_model(kind="nearest", name="model"):    return Model(kind, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  examples = new_dataset("examples")examples.add("Please explain gravity", "question")examples.add("Could you explain magnets", "question")examples.add("Why does ice melt", "question")examples.add("How does a battery work", "question")examples.add("Where can I find the library", "question")examples.add("Can you describe a volcano", "question")examples.add("What causes thunder", "question")examples.add("How do seeds grow", "question")examples.add("The parcel arrived safely", "statement")examples.add("My parcel arrived yesterday", "statement")examples.add("Dinner tastes delicious", "statement")examples.add("Our team won today", "statement")examples.add("The blue door is closed", "statement")examples.add("Music played softly", "statement")examples.add("The dog slept peacefully", "statement")examples.add("We painted the fence", "statement")judge = new_model("words", "judge")judge.train(examples)message = input()label = judge.predict(message)sure = judge.confidence(message)print("Input", message)if (sure < 60):    print("I am not sure. Please rephrase.")else:    if (label == "question"):        print("That looks like a question.")    else:        print("That looks like a statement.")print("Raw label", label)print("Raw score", sure)

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in