Prev: 6E8B Up: Map Next: 6F30
6E97: Turn the next word of INPUT_LINE into a token
Used by the routine at START.
A token is two bytes: the word's class in the top nibble -- bits 5-6 of its first two dictionary bytes, read together -- and its twelve-bit dictionary offset below that. A synonym comes out as the word it stands for, so GET SWORD is TAKE SWORD by the time anything reads it. $C0 ends the line; $D0 is a word not in the dictionary, and the main loop prints the complaint and never calls the parser.
A typed word may be shortened or, within limits, lengthened. A candidate is taken if the two agree over the shorter length and the typed word is the shorter -- an abbreviation; if the typed word is the longer, only when the entry has at least four letters and no later candidate fits too. Tried: EXAM gives EXAMINE, INV gives INVENTORY, SWORDS gives SWORD, INT gives INTO because IN is too short to stretch, and EXAMINING is not a word at all, since its seventh letter disagrees.
Watched on real sentences: VICIOUSLY ATTACK THE TROLL WITH THE SWORD comes out as adverb, verb, article, noun, preposition, article, noun, end; TAKE THE MAP AND THE KEY puts AND in class $A; and a closing quote gets a full stop token inserted before it by the main loop, so what is said to a character ends as a sentence.
Output
BC The token
TOKENISE 6E97 PUSH DE
TOKENISE_0 6E98 LD A,(HL) Skip spaces
6E99 INC HL
6E9A CP $20
6E9C JR Z,TOKENISE_0
6E9E DEC HL
6E9F LD (VARIABLES),HL Remember where the word starts, for the echo of an unknown one
6EA2 CP $0D The end of the line: $C0
6EA4 JR Z,TOKENISE_5
6EA6 CALL PUNCTUATION_TOKEN A full stop, comma or quote is a token by itself
6EA9 JR Z,TOKENISE_7
6EAB CALL MATCH_WORD Start on the dictionary bucket for its first letter; none, and it is unknown
6EAE JR NZ,TOKENISE_3
6EB0 PUSH HL Does this candidate agree with what was typed?
TOKENISE_1 6EB1 CALL LETTERS_AGREE
6EB4 JR Z,TOKENISE_4
TOKENISE_2 6EB6 CALL TRY_ENTRY No: try the next in the bucket, until it runs out
6EB9 JR Z,TOKENISE_1
6EBB POP HL
TOKENISE_3 6EBC LD A,$D0 Not in the dictionary: $D0
6EBE JR TOKENISE_6
TOKENISE_4 6EC0 LD A,(TYPED_LENGTH) The typed word no longer than the entry: an abbreviation, take it
6EC3 LD B,A
6EC4 LD A,(ENTRY_LENGTH)
6EC7 CP B
6EC8 JR NC,TOKENISE_9
6ECA CP $04 Longer, and the entry under four letters: not this one
6ECC JR C,TOKENISE_2
6ECE PUSH IX Longer, and the entry four or more: take it unless the next candidate agrees too
6ED0 CALL TRY_SAME_ENTRY
6ED3 JR NZ,TOKENISE_8
6ED5 CALL LETTERS_AGREE
6ED8 JR NZ,TOKENISE_8
6EDA POP IX
6EDC JR TOKENISE_2
TOKENISE_5 6EDE LD A,$C0 No word: end of line, or unknown, or punctuation
TOKENISE_6 6EE0 LD BC,$0000
TOKENISE_7 6EE3 POP DE B = class and top of the offset, C the low byte; A = the class
6EE4 LD D,A
6EE5 ADD A,B
6EE6 LD B,A
6EE7 LD A,D
6EE8 RET
TOKENISE_8 6EE9 POP IX
TOKENISE_9 6EEB LD IX,(DICTIONARY_ENTRY) Walk to the end of the chosen entry...
6EEF PUSH IX
6EF1 XOR A
TOKENISE_10 6EF2 INC IX
6EF4 INC A
6EF5 BIT 7,(IX-$01)
6EF9 JR Z,TOKENISE_10
6EFB CP $02 ...by PRINT_WORD's rule for where a word ends
6EFD JR Z,TOKENISE_10
6EFF CP $03
6F01 JR NZ,TOKENISE_11
6F03 BIT 7,(IX-$02)
6F07 JR NZ,TOKENISE_10
TOKENISE_11 6F09 BIT 6,(IX-$01) A synonym: its link, turned into an address, replaces it
6F0D JR Z,TOKENISE_12
6F0F LD L,(IX+$00)
6F12 LD H,(IX+$01)
6F15 LD DE,WORD_INDEX
6F18 ADD HL,DE
6F19 EX (SP),HL
TOKENISE_12 6F1A POP HL The class: bits 5-6 of the first byte above bits 5-6 of the second
6F1B LD A,(HL)
6F1C RLCA
6F1D AND $C0
6F1F LD B,A
6F20 INC HL
6F21 LD A,(HL)
6F22 RRCA
6F23 AND $30
6F25 ADD A,B
6F26 DEC HL And the entry's offset from WORD_INDEX
6F27 LD DE,$A000
6F2A ADD HL,DE
6F2B PUSH HL
6F2C POP BC
6F2D POP HL
6F2E JR TOKENISE_7
Prev: 6E8B Up: Map Next: 6F30