![]() |
Routines |
| Prev: 6E8B | Up: Map | Next: 6F30 |
|
Used by the routine at START.
A token is two bytes: the word's class in the top nibble -- bits 5-6 of its first two dictionary bytes, read together -- and its twelve-bit dictionary offset below that. A synonym comes out as the word it stands for, so GET SWORD is TAKE SWORD by the time anything reads it. $C0 ends the line; $D0 is a word not in the dictionary, and the main loop prints the complaint and never calls the parser.
A typed word may be shortened or, within limits, lengthened. A candidate is taken if the two agree over the shorter length and the typed word is the shorter -- an abbreviation; if the typed word is the longer, only when the entry has at least four letters and no later candidate fits too. Tried: EXAM gives EXAMINE, INV gives INVENTORY, SWORDS gives SWORD, INT gives INTO because IN is too short to stretch, and EXAMINING is not a word at all, since its seventh letter disagrees.
Watched on real sentences: VICIOUSLY ATTACK THE TROLL WITH THE SWORD comes out as adverb, verb, article, noun, preposition, article, noun, end; TAKE THE MAP AND THE KEY puts AND in class $A; and a closing quote gets a full stop token inserted before it by the main loop, so what is said to a character ends as a sentence.
|
||||||||
| TOKENISE | 6E97 | PUSH DE | ||||||
| TOKENISE_0 | 6E98 | LD A,(HL) | Skip spaces | |||||
| 6E99 | INC HL | |||||||
| 6E9A | CP $20 | |||||||
| 6E9C | JR Z,TOKENISE_0 | |||||||
| 6E9E | DEC HL | |||||||
| 6E9F | LD (VARIABLES),HL | Remember where the word starts, for the echo of an unknown one | ||||||
| 6EA2 | CP $0D | The end of the line: $C0 | ||||||
| 6EA4 | JR Z,TOKENISE_5 | |||||||
| 6EA6 | CALL PUNCTUATION_TOKEN | A full stop, comma or quote is a token by itself | ||||||
| 6EA9 | JR Z,TOKENISE_7 | |||||||
| 6EAB | CALL MATCH_WORD | Start on the dictionary bucket for its first letter; none, and it is unknown | ||||||
| 6EAE | JR NZ,TOKENISE_3 | |||||||
| 6EB0 | PUSH HL | Does this candidate agree with what was typed? | ||||||
| TOKENISE_1 | 6EB1 | CALL LETTERS_AGREE | ||||||
| 6EB4 | JR Z,TOKENISE_4 | |||||||
| TOKENISE_2 | 6EB6 | CALL TRY_ENTRY | No: try the next in the bucket, until it runs out | |||||
| 6EB9 | JR Z,TOKENISE_1 | |||||||
| 6EBB | POP HL | |||||||
| TOKENISE_3 | 6EBC | LD A,$D0 | Not in the dictionary: $D0 | |||||
| 6EBE | JR TOKENISE_6 | |||||||
| TOKENISE_4 | 6EC0 | LD A,(TYPED_LENGTH) | The typed word no longer than the entry: an abbreviation, take it | |||||
| 6EC3 | LD B,A | |||||||
| 6EC4 | LD A,(ENTRY_LENGTH) | |||||||
| 6EC7 | CP B | |||||||
| 6EC8 | JR NC,TOKENISE_9 | |||||||
| 6ECA | CP $04 | Longer, and the entry under four letters: not this one | ||||||
| 6ECC | JR C,TOKENISE_2 | |||||||
| 6ECE | PUSH IX | Longer, and the entry four or more: take it unless the next candidate agrees too | ||||||
| 6ED0 | CALL TRY_SAME_ENTRY | |||||||
| 6ED3 | JR NZ,TOKENISE_8 | |||||||
| 6ED5 | CALL LETTERS_AGREE | |||||||
| 6ED8 | JR NZ,TOKENISE_8 | |||||||
| 6EDA | POP IX | |||||||
| 6EDC | JR TOKENISE_2 | |||||||
| TOKENISE_5 | 6EDE | LD A,$C0 | No word: end of line, or unknown, or punctuation | |||||
| TOKENISE_6 | 6EE0 | LD BC,$0000 | ||||||
| TOKENISE_7 | 6EE3 | POP DE | B = class and top of the offset, C the low byte; A = the class | |||||
| 6EE4 | LD D,A | |||||||
| 6EE5 | ADD A,B | |||||||
| 6EE6 | LD B,A | |||||||
| 6EE7 | LD A,D | |||||||
| 6EE8 | RET | |||||||
| TOKENISE_8 | 6EE9 | POP IX | ||||||
| TOKENISE_9 | 6EEB | LD IX,(DICTIONARY_ENTRY) | Walk to the end of the chosen entry... | |||||
| 6EEF | PUSH IX | |||||||
| 6EF1 | XOR A | |||||||
| TOKENISE_10 | 6EF2 | INC IX | ||||||
| 6EF4 | INC A | |||||||
| 6EF5 | BIT 7,(IX-$01) | |||||||
| 6EF9 | JR Z,TOKENISE_10 | |||||||
| 6EFB | CP $02 | ...by PRINT_WORD's rule for where a word ends | ||||||
| 6EFD | JR Z,TOKENISE_10 | |||||||
| 6EFF | CP $03 | |||||||
| 6F01 | JR NZ,TOKENISE_11 | |||||||
| 6F03 | BIT 7,(IX-$02) | |||||||
| 6F07 | JR NZ,TOKENISE_10 | |||||||
| TOKENISE_11 | 6F09 | BIT 6,(IX-$01) | A synonym: its link, turned into an address, replaces it | |||||
| 6F0D | JR Z,TOKENISE_12 | |||||||
| 6F0F | LD L,(IX+$00) | |||||||
| 6F12 | LD H,(IX+$01) | |||||||
| 6F15 | LD DE,WORD_INDEX | |||||||
| 6F18 | ADD HL,DE | |||||||
| 6F19 | EX (SP),HL | |||||||
| TOKENISE_12 | 6F1A | POP HL | The class: bits 5-6 of the first byte above bits 5-6 of the second | |||||
| 6F1B | LD A,(HL) | |||||||
| 6F1C | RLCA | |||||||
| 6F1D | AND $C0 | |||||||
| 6F1F | LD B,A | |||||||
| 6F20 | INC HL | |||||||
| 6F21 | LD A,(HL) | |||||||
| 6F22 | RRCA | |||||||
| 6F23 | AND $30 | |||||||
| 6F25 | ADD A,B | |||||||
| 6F26 | DEC HL | And the entry's offset from WORD_INDEX | ||||||
| 6F27 | LD DE,$A000 | |||||||
| 6F2A | ADD HL,DE | |||||||
| 6F2B | PUSH HL | |||||||
| 6F2C | POP BC | |||||||
| 6F2D | POP HL | |||||||
| 6F2E | JR TOKENISE_7 | |||||||
| Prev: 6E8B | Up: Map | Next: 6F30 |