A record is one set of related items, and each item is a field. A file may store records as lines of text.
The program has to cut each line back into fields without moving a boundary. It then converts numbers from text if it needs to calculate with them.
This lesson sits in searching, sorting and files. The data here is fictional, and the next lesson, handling end-of-file, shows how to read many such lines.
How can a line be split into fields?
There are two common designs.
Delimited fields. A marker such as a comma separates the fields. The line Aina,10B,72 has three fields because it has two commas.
Fixed-width fields. Each field always uses the same number of characters, padded with spaces or zeros. The program counts characters instead of looking for a marker.
Worked example: delimited
The line is Aina,10B,72, holding a name, a class and a score. In Python:
line = "Aina,10B,72"
fields = line.split(",")
name = fields[0]
cls = fields[1]
score = int(fields[2])
print(name, cls, score + 5)
line.split(",") gives ["Aina", "10B", "72"]. Python counts positions from 0, so the name is fields[0] and the score is fields[2].
The score arrives as the text "72", so int converts it to the number 72. Then score + 5 is 77, and the output is Aina 10B 77.
Cambridge-style pseudocode counts differently and has its own string functions, so check the current pseudocode guide for your exam year. The reasoning about boundaries is the same.
Worked example: fixed-width
The line is Aina 10B072 with a name field of 6 characters, a class field of 3 and a score field of 3.
| Field | Characters | Value |
|---|---|---|
| Name | 1 to 6 | Aina (the name plus two spaces) |
| Class | 7 to 9 | 10B |
| Score | 10 to 12 | 072, which is the number 72 |
Check the count: Aina is 4 characters plus 2 spaces makes 6, then 10B makes 9, then 072 makes 12. In pseudocode, MID(Line, 7, 3) returns the 3 characters starting at position 7, which is 10B. The score field MID(Line, 10, 3) returns 072.
The mistake to watch for
A student reads two scores from a file as text and compares them:
Is
"9"greater than"72"?
As text, the comparison goes character by character: "9" against "7". The character 9 comes after 7, so the text "9" is judged greater. As numbers, 9 is clearly less than 72.
The correction is to convert the field to a number before comparing or adding. When you trace, write the data type next to each field. A second boundary slip is a name such as Tan, Mei in a comma file: the split gives four fields instead of three, so the class lands in the score position.
Check yourself
1. Split Zara,11A,64 at the commas. What is the third field, and what is it plus 6?
Show answer
Fields: Zara, 11A, 64. The third field is "64", converted to the number 64. Then 64 + 6 = 70.
2. A fixed-width line Omar 09C055 uses widths 6, 3 and 3. State each field.
Show answer
Characters 1 to 6 are Omar . Characters 7 to 9 are 09C. Characters 10 to 12 are 055, which is the number 55.
3. The line Lee, Ann,10C,81 is split at commas. How many fields result, and what has gone wrong?
Show answer
Three commas give four fields: Lee, Ann, 10C, 81. The delimiter appears inside the name, so every later field shifts one place.
Where this leads next
With fields under control, move on to reading a whole file until the end. The safe Python reasoning sandbox lets you try a split like the one above and see each field printed.
If records and conversions still feel slippery on new lines, our teachers can build extra examples with you in online one-to-one Computer Science tuition.