[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Integer Basic Tokenization



G'day,

I'm currently writing a utility to manage my Apple ][ disk images. As part
of this I'm going through a lot of my old personal disks that I created when
I was first learning computers in the early 80s, including a lot of BASIC
programs I typed in from magazines and so on. (Remember when you used to be
able to do that? Computer Magazines came with "Special 16 page pull-out
program supplements", not DVDs full of code ready to go.)

Anyway, I digress...

I've sucessfully written a program to take an APPLESOFT program and
de-tokenize it back into readable text, and it works an absolute treat. But
INTEGER BASIC programs are proving a little bit more difficult, and I was
wondering if somebody out there could shed some light on the process it
uses.

So far I have the structure being like this:

1 Byte: Length of Line
2 Bytes: Line Number (Lo/Hi Order)
? Bytes Tokenized Program
Last Byte: 01 (End of line token)

So, pulling apart some code I get:

105 PRINT "[CTRL-D]BLOAD BOWLING.OBJ"
Hex: 19 69 00 61 28 84 C2 CC CF C1 C4 A0 C2 CF D7 CC C9 CE C7 AE CD C2 CA 29
01

19 = Line is 25 Bytes long
69 00 = Line 105
61 = PRINT token
28 = Quote Token
84 = Ctrl-D Character
C2 CC CF C1 C4 A0 C2 CF D7 CC C9 CE C7 AE CD C2 CA = ASCII String (Hi bit
Set)
29 = Quote Token (Different for closing quote... interesting.)
01 = End of Line

That's not too bad and is quite similar to Applesoft except that things like
quotes are tokenized and plain text has the high bit set. But once numbers
start appearing in the code, things get really messy. INTEGER appears to
encode all numbers too, whereas APPLESOFT just has them as plain text. So we
get:

108  LOMEM: 5000
Hex: 08 6C 00 11 B5 88 13 01

08 = Line is 8 bytes long
6C 00 = Line 108
11 = LOMEM Token
B5 = Colon Token?? Or pointer that the next bytes are a number?
88 13 = 5000 (Stored in Lo/Hi Order)
01 = End of line

110 POKE 808,0  : POKE 809,12
Hex: 15 6E 00 64 B8 28 03 65 B0 00 00 03 64 B8 29 03 65 B1 0C 00 01

15 = Line is 21 bytes long
6E 00 = Line 110
64 = POKE Token
B8 = ??
28 03 = 808 (Lo/Hi)
65  = Comma token?
B0 = ??
00 00 = 0 (Lo,Hi)
03 = Colon Token
64 = POKE Token
B8 = ??
29 03 = 809 (Lo/Hi)
65 = Comma token
B1 = ??
0C 00 = 12 (Lo/Hi)
01 = End of Line

Particularly confusing in this case is that B0 appears after the comma token
in the first poke, but B1 appears after the comma token in the second poke
statement. It would appear that the B? character matches the first digit of
the number that follows it, but that seems a bit weird to me, and certainly
isn't an infallible coding.

Can anyone help? (And just a list of tokens would be helpful!)


Regards,


Michael