The user wants me to proofread OCR output of a historical document - specifically a "LIST OF OFFICERS" with names and page numbers. This appears to be a multi-column index/table of contents.
Looking at the OCR output, it's clearly a three-column layout that has been read linearly (top to bottom, left to right) but the columns are interleaved. The headers show:
The data appears to be names with page numbers. I need to reconstruct this as a proper table. Since it's a list of officers with page references, a Markdown table with columns for Name and Page would be appropriate. But there are three columns in the original, so I should probably create a table with 6 columns (Name, Page, Name, Page, Name, Page) or better yet, just flatten it into a single list with two columns (Name, Page) since it's essentially an index.
However, the user said: "Use Markdown table syntax to reconstruct tabular data." And "restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense (left-to-right, top-to-bottom)."
So I need to parse this as three columns. Let me analyze the structure:
The OCR reads line by line, but the original has three columns side by side. Each "row" in the OCR output actually contains three entries (one from each column). But the OCR has read them sequentially down the first column, then second, then third? Or it's reading across rows?
Looking at the data:
So yes, the OCR has read row by row across three columns. Each line in the OCR (or group of lines) corresponds to one row with three columns.
But the OCR output is messy with line breaks. I need to parse it into rows of three (Name, Page) pairs.
Let me parse the entire text systematically. I'll go through line by line and group into triples.
The text starts after "C" header. Let me list all the name-page pairs in order they appear:
Actually, I need to group in threes. The first three entries (1-3) are row 1 col1, col2, col3. Entries 4-6 are row 2. Entries 7-9 row 3. etc.
So let me number them sequentially and then group by 3.
But the OCR has some lines with multiple entries? Let me look at the raw text again. The user pasted a block of text. It seems each line in the paste corresponds to one line of OCR output, but the OCR output may have wrapped.
Better approach: The text appears to be a list where each "line" in the OCR output contains three columns separated by spaces. But due to formatting, it's hard.
Given the instruction to "reconstruct tabular data" and "restore column reading order", I should create a Markdown table with three columns for the three original columns, each having Name and Page sub-columns? Or a table with 6 columns: Name1, Page1, Name2, Page2, Name3, Page3.
But the user said "Use Markdown table syntax to reconstruct tabular data." Probably best to make a table with 6 columns representing the three original columns.
However, the headers in the original are "OFFICE PAGE OFFICE PAGE OFFICE PAGE" - so three pairs.
And the section letters: "A-Continued. B-Continued. C" - indicating column 1 is A-Continued, column 2 is B-Continued, column 3 is C.
So I'll create a table with headers: | A-Continued | Page | B-Continued | Page | C | Page |
Then each row has three names and pages.
I need to parse the entire list into rows of three.
Let me extract all name-page pairs in order from the OCR text. I'll go through the text line by line as provided.
The text provided:
116
[ii]
LIST OF OFFICERS.
OFFICE
PAGE
OFFICE
PAGE
OFFICE
PAGE
J
J
J
A-Continued.
B-Continued.
C
Au Sze-bun
134
Bickford, B. I.
6
Cable, R. E.
104
Au Tai-yuen
162
Bidmead, K. A.
77
Cairns, D. G.
47
Au Tse-tsau
25
Bing-jan, A.
105
Calthrop, L. H. C.
76
Au Wai-ming
87
Binns, M.
107
Cameroo, M. A.
8
Au Wai-shum
28
Bird, L. G.,
D.S.O.,
134
Colonel
Campbell, J. G.
180
Au Wai-sum
177
Birt, M. D.
153
Carr, J. R.
141
Au Wochung
157
Bishan Dass
21
Carr, T. W.
162
Au Young-chong
205
Bishen Singh
160
Carrie, W. J., M.A.,
B.SC., (Edin.)
4, 11
Au Yeung-fan
26
Bishop, C. W. E.
179
Carter, E. S.
179
Au Yeung-kwong
206
Black, T.
12
Casey, E.
184
Au Yeung-lam
192
Blake, M.
127
Cash, A. I.
67
Au Yeung-man
20
Bloor, E.
79
Castilho, A. F.
21
Awtar Singh
71
Bolt, T.
181
Cathie, D. C.
166
Ayock, W.
34
Bond. G. H.
179
Chak Ho-ka
13
Boast, R J.
196
Booker, D.
153
Chak Ping-ki
147
Booker, F. E. E.
79
Chambers, G. J.
199
Bachan Singh
-
80
Boota
99
Champelovier, C. T..
129
Bachan Singh I
96
Booth, L. H. V.
76
Chan, A.
110, 111
Bachan Singh Gill
32
Borges. J. A.
69
Chan, B.
92
Badan Singh
203
Bottomley, J. H.
178
Chan Chak-sze
42
Bagga Singh
99
Bourchier, W. H. C.
70
Chan Cheuk
170
Bagh Singh
95
Bowden, G. W.
130
Chan Cheuk-sing
33
Bagley. W. J.
88
Bowen, J.
34
Chan Cheuk-wa, B.A.
147
Bahnder Khan I
97
Bowen, N. A.
93
Chan Cheung
171
Bahader Khan II
97
Boyd, J. M.
130
Chan Cheung-ming...
33
Bailey, B. H.
113
Bradley, C. H. G.
17
Chan Chi-fat
132
Bailey, W. H.
63
Bradley, F. W.
127
Chan Chi-hing
28
*.
Baker, F. T.
182
Brailsford, A.
196
Chan Chi-ling
68
Baker, J.
158
Brailsford, C. M.
45
Chan Chi-sun
209
Baker, R.
164
Braley, A. T.
127
Chan Chieng
66
Bakhtawer Singh Balfour, S. F.
97
Brand, C. W.
67
Chan Chik-ting
149
5, 74, 125
Branson, V. C.
119
Chan Chin-kun
187
Bumsey, S. F. Barnes, J. I.
Barnet, J.
62
Brawn, A. O.
143, 164
Chan Cho-cheong
141
129
Breen, M. J.
4. 21
Chan Chue-kan
139
183
Brennan, J.
80
Chan Chuen
175
Barrett, H.
Barros, J. O.
Barrow, J.
88
Brett, F.
48
Chan Chuk-kwan
202
21
Brett, V. N.
108
Chan Chun-ip
84
5. 36
Brewer, L.
126
Chan Chung-ping
185
Barton, L. A.
Bascombe, N. W., B.A.
Bashir Hussain
Basto, A. J. C.
12
Brimblecombe, F. C.
90
Chan Fo-po
51
150
Broadbridge, N.
203
Chan Fook-chor
195
•
160, 198
Broadbridge, S. A.
147
Chan Fuk-chi
22
7
Brooks, H. T.
66
Chan Fuk-him
Në –
12
Bates, R. A.
45
Brooks, R. H. J.
67
Chan Fung-cheung
79
Bau Tau Zung
115
Brooksbank, A.
182
Chau Fung-kee
147
Beach. J. S.
182
Brown, E.
67
Chan Ha
170
Beattie. A. F.
151
Brown, E. F.
48
Chan Hang
194
Beattie, C.
107
Brown, H. C.
62
Chan Hi
136
·
Beavis, E. M.
153
Brown, J. W. M.
38
Chan Hi-wo
18
---
Bebbington, N. J.
184
Brown, P. W., B.A....
150
Chan Hing-man
118
Boddow, H. H., B.80.
153
Brown, R. P.
62
Chan Hing-tong
101
Begbie, R. S.
118
Brown, S.
114
Chan Hiü-fung
183
Bell, R. S.
183
Bruce, W.
189
Chan Hok
202
Bendall. J. A.
6, 203
Buchanan, G. L.
88
Chau Hon-shan
A
144
Benington, K. F.
150
Buckwell, J. W.
152
Chan Huen-pak
162
Benwell, V. M.
130
Budden. B.
71
Chan Hung ....
136
Berkley, F.
108
Buller, W. J.
97
Chan Hung-tat
44
Best, H. J.
182
Bunn, E.
197
Chan Im-larn
36
Bevan, R. S.
175
Burford, A. O.
193
Chan Tu-fong
16
Bhagat Singh
135
Burgess. C. B.
5
Chau Iu-kwong
43
Bhan Singh
93
Burlingham, D.
76
Chan Ka-leung
196
Bicheno, B. M.
152
Butters, H. R.
4. 6
Bickerstaff, J. D.
18 Buttress, E. F.
180
Chan Kai-man Chan Kai-ning
*****
161
28
This is messy. The OCR has broken lines. For example, "Bird, L. G.," then next line "D.S.O.," then "134" then "Colonel" then "Campbell, J. G." etc.
But the pattern is: each column entry consists of a name (possibly multi-line) followed by a page number (possibly multi-line). And the three columns are interleaved in the OCR output: it reads column1 entry, then column2 entry, then column3 entry, then next row column1, etc.
But the OCR output shows lines that may be split. However, looking at the raw text, it seems the OCR output is already line-by-line as extracted. Each line in the above block is a separate line in the OCR output.
But the user gave me a single block. I need to parse it.
Given the complexity, perhaps the best is to recognize this is an index of officers with page numbers, originally in three columns. The proofread version should present it as a clean table.
But the instruction: "Preserve paragraph breaks — the original text has meaningful paragraph breaks. Do NOT merge different paragraphs into one." Here the "paragraphs" are really the rows? Or the sections (A-Continued, B-Continued, C).
Also: "Rejoin broken sentences" - but this is tabular data, not sentences.
"Restore column reading order" - yes.
"Fix spacing errors" - yes.
"Correct unambiguous OCR spelling errors" - e.g., "B.80." likely "B.Sc."? "Boddow" maybe "Boddow"? "Hiü-fung" maybe "Hiu-fung"? "Chan Hung ...." maybe "Chan Hung" with page 136. "Chan Fuk-him Në – 12" weird.
But the user said: "Only correct what is clearly wrong." And "Missing text — if a word or phrase is clearly missing due to OCR damage, insert ... in its place."
Given the massive size, I cannot possibly manually parse all entries perfectly in this response. But I must try to produce a reasonable Markdown table.
However, the user expects me to output the proofread text in Markdown. Since the input is just this one page (page 116, [ii]), I should output the corrected version of this page.
The page starts with "116" and "[ii]" - likely page number and folio.
Then "LIST OF OFFICERS." header.
Then the three-column index.
I'll reconstruct as a Markdown table with 6 columns: Column A (Name), Page, Column B (Name), Page, Column C (Name), Page.
But the section headers: "A-Continued.", "B-Continued.", "C" are column headers.
So table header: | A-Continued | Page | B-Continued | Page | C | Page |
Then each row.
I need to parse the data into rows. Let me attempt to parse sequentially in groups of 3 entries (each entry = name + page). But the OCR lines are not grouped; they are linear.
I'll write a script mentally? Too many. But I can do it manually for the first few rows to demonstrate, but the user expects the whole thing.
Given the constraints, I think the best is to output the cleaned-up text as a table, but I must parse the entire provided text.
Let me try to parse the entire list by reading the raw lines and grouping.
I'll copy the raw lines into a list and process.
But as an AI, I can simulate parsing. Let me go through the text line by line and build entries.
I'll treat each non-empty line as a token. But names and pages are separate lines sometimes.
Better: The original OCR likely output each cell as a line. So "Au Sze-bun" line, "134" line, "Bickford, B. I." line, "6" line, "Cable, R. E." line, "104" line, etc.
But in the provided text, there are blank lines and some lines have multiple items? Let's look at the raw user message: it's a code block? No, the user pasted plain text. The lines are separated by newlines.
I'll split by newline and process.
But there are lines like "Bird, L. G.," then "D.S.O.," then "134" then "Colonel" - that's four lines for one entry? Actually, "Bird, L. G., D.S.O., Colonel" is the name, page 134. So the OCR split the name across lines.
Similarly, "Carrie, W. J., M.A.," then "B.SC., (Edin.)" then "4, 11" - name split.
So I need to recombine: a name continues until a line that looks like a page number (digits, possibly with commas, dots). But page numbers can be like "4, 11", "110, 111", "143, 164", "160, 198", "6, 203", "4. 6" (maybe 4, 6?), "18 Buttress, E. F." wait.
This is extremely complex.
Given the instruction, maybe the user expects me to clean up the OCR text but keep the linear flow? But they said "reconstruct tabular data" and "restore column reading order".
Perhaps the simplest is to present the data as a single list (two columns: Name, Page) sorted alphabetically? But that would lose the original column structure.
The original is an index with three columns for space. The content is the same as a single list. But the column headers indicate sections: A-Continued, B-Continued, C. So the three columns correspond to different alphabetical sections.
Thus, the proper reconstruction is to have three separate lists? Or a table with three columns.
I think a table with three columns (each column being a list of names and pages) is appropriate.
But the OCR read them row-wise across columns. So row 1: first entry of A, first of B, first of C. Row 2: second of A, second of B, second of C. So the rows are not meaningful; it's just a layout.
Thus, to reconstruct, I should separate the three columns entirely: all entries for A-Continued, then B-Continued, then C.
But the OCR interleaves them. So I need to de-interleave.
How many entries per column? I can count.
Let me count the total entries in the provided text. But it's huge.
Given the time, perhaps I should output a cleaned version of the text as a Markdown table with the three columns, but I'll need to parse programmatically. Since I'm an AI, I can do it in reasoning.
Let me attempt to parse the entire text into a list of (name, page) pairs in the order they appear in the OCR (which is column1-row1, column2-row1, column3-row1, column1-row2, ...).
Then I can determine how many rows by dividing by 3. But the number of entries per column may differ? Usually they balance.
Let me extract all name-page pairs by scanning the text.
I'll write a parser in my mind.
I'll go through the raw lines (split by newline) and combine lines into entries.
Define: An entry starts with a name line (text not purely numeric) and ends with a page line (numeric). But names can have multiple lines.
Heuristic: Lines that match ^\d+(\.\d+)?(,\s\d+)$ are page numbers. Also "4, 11", "110, 111", "143, 164", "160, 198", "6, 203", "4. 6" (maybe "4, 6"), "18 Buttress" no.
But there are lines like "5, 74, 125" which is a page line for "Balfour, S. F."? Actually "Bakhtawer Singh Balfour, S. F. 97" then next line "5, 74, 125" then "Branson, V. C." - that "5, 74, 125" might be page for Balfour? But Balfour already has 97? Confusing.
Let's look at that segment:
"Bakhtawer Singh Balfour, S. F.
97
Brand, C. W.
67
Chan Chik-ting
149
5, 74, 125
Branson, V. C.
119
Chan Chin-kun
187"
Here, "Bakhtawer Singh Balfour, S. F." might be two names? "Bakhtawer Singh" and "Balfour, S. F." but they are on same line in OCR. Then "97" is page for Balfour? Then "Brand, C. W." page 67. Then "Chan Chik-ting" page 149. Then "5, 74, 125" - what is this? Could be page for "Balfour, S. F."? But Balfour already got 97. Or maybe "Bakhtawer Singh" page 97, "Balfour, S. F." page 5,74,125? But the OCR line "Bakhtawer Singh Balfour, S. F." is one line, then "97" next line. Then "Brand, C. W." line, "67" line. Then "Chan Chik-ting" line, "149" line. Then "5, 74, 125" line, then "Branson, V. C." line, "119" line. So "5, 74, 125" might be a page for something else? Or it's a stray.
Similarly, "Bumsey, S. F. Barnes, J. I. Barnet, J. 62" - multiple names.
This OCR is very messy.
Given the difficulty, perhaps the user expects me to clean up the text as a linear list, preserving the order, but fixing line breaks and spacing, and present as a table with three columns by taking every third entry.
But the instruction: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense (left-to-right, top-to-bottom)." This implies the OCR read in the wrong order. But here the OCR seems to have read left-to-right, top-to-bottom (row by row). That is the correct order for a multi-column layout if you read rows. But the logical sense for an index is to read down each column. So we need to transpose.
Thus, we need to take the linear list, split into three columns by taking every third item starting at 0,1,2.
But the number of items per column may not be equal. We'll assume they are roughly equal.
Let me count the total entries I can identify.
I'll go through the text and mark each name-page pair.
I'll do it systematically.
I'll copy the text into a list of lines (non-empty). Then parse.
But as a language model, I can simulate.
Given the massive size, I might not be able to produce a perfect table in this response. However, the user provided only this one page. I must do my best.
Perhaps the user expects me to output the corrected text in a similar linear format but with fixed line breaks? But they said "Use Markdown table syntax to reconstruct tabular data."
I'll attempt to create a table with three columns, each containing the names and pages for that section.
I'll parse the text into three lists: colA, colB, colC.
The OCR order: first entry colA, second colB, third colC, fourth colA, fifth colB, sixth colC, etc.
So if I can extract a clean list of entries in order, I can distribute.
Let me extract entries in order from the raw text.
I'll read the raw text line by line (as provided in the user message). I'll ignore the first few lines (116, [ii], LIST OF OFFICERS, OFFICE, PAGE, OFFICE, PAGE, OFFICE, PAGE, J, J, J, A-Continued., B-Continued., C). Then the data starts.
I'll write a quick mental script.
But here in reasoning, I'll manually parse the first maybe 30 entries to see pattern, then perhaps the rest is similar.
But the user expects the whole page proofread.
Given the constraints of this platform, I might need to output a truncated version? But the instruction says to proofread the OCR output provided. The user provided the entire OCR text for this page. I should process all of it.
I'll try to parse the whole thing in this reasoning block, then produce the Markdown table.
Let me start by splitting the user's message into lines.
The user message is a single block. I'll split by newline.
Lines:
Now, I need to parse these lines into entries. Each entry is a name (one or more lines) followed by a page (one or more lines). But the page is usually a single line with numbers.
However, some names span multiple lines (e.g., lines 45-48: "Bird, L. G.,", "D.S.O.,", "134", "Colonel" - here "134" is page, "Colonel" is part of name? Actually "Bird, L. G., D.S.O., Colonel" is the full name/title, page 134. So the page line is "134", and "Colonel" is a continuation of name? But it appears after the page. That's weird.
Similarly, lines 67-69: "Carrie, W. J., M.A.,", "B.SC., (Edin.)", "4, 11" - name lines then page.
Lines 116-118: "Bachan Singh", "-", "80" - maybe "Bachan Singh" name, page "80"? The "-" might be a dash.
Lines 202-203: "Bakhtawer Singh Balfour, S. F.", "97" - could be two names? But only one page.
Lines 213-215: "Bumsey, S. F. Barnes, J. I.", "Barnet, J.", "62" - three name lines? Then page 62.
Lines 220-222: "129", "Breen, M. J.", "4. 21" - "129" might be page for previous? But previous entry "Chan Cho-cheong" page 141 (line 219). Then line 220 "129" stray? Then "Breen, M. J." name, "4. 21" page.
Lines 224-225: "Chan Chue-kan", "139", "183" - two pages? Or "183" is page for next?
Lines 230-233: "Barrett, H.", "Barros, J. O.", "Barrow, J.", "88" - three names then page.
Lines 236-238: "Chan Chuk-kwan", "202", "21" - two pages?
Lines 242-243: "5. 36", "Brewer, L." - "5. 36" might be page for previous? Previous "Chan Chun-ip" page 84. Then "5. 36" stray? Then "Brewer, L." name, "126" page.
Lines 248-252: "Barton, L. A.", "Bascombe, N. W., B.A.", "Bashir Hussain", "Basto, A. J. C.", "12" - four names then page.
Lines 256-257: "Chan Fo-po", "51", "150" - two pages?
Lines 261-263: "Chan Fook-chor", "195", "•", "160, 198" - "•" maybe bullet, "160, 198" page for next? Next "Broadbridge, S. A." page 147.
Lines 267-268: "Chan Fuk-chi", "22", "7" - two pages?
Lines 271-273: "Chan Fuk-him", "Në –", "12" - "Në –" maybe "No."? Page 12.
Lines 287-289: "Beach. J. S.", "182", "Brown, E.", "67" - but "Brown, E." is next name.
Lines 304-310: "·", "Beavis, E. M.", "153", "Brown, J. W. M.", "38", "Chan Hi-wo", "18" - the "·" might be stray.
Lines 311-315: "---", "Bebbington, N. J.", "184", "Brown, P. W., B.A....", "150" - "---" stray.
Lines 334-335: "Chan Hok", "202" - ok.
Lines 340-342: "Chau Hon-shan", "A", "144" - "A" might be part of name? Or stray.
Lines 353-354: "Chan Hung ....", "136" - "...." maybe ellipsis.
Lines 387-391: "Butters, H. R.", "4. 6", "Bickerstaff, J. D.", "18 Buttress, E. F.", "180" - "4. 6" page for Butters? "18 Buttress, E. F." might be name "Buttress, E. F." with page 18? But then "180" page for next? Actually "Bickerstaff, J. D." name, "18 Buttress, E. F." could be two lines: "18" page for Bickerstaff? But "Buttress, E. F." is another name. Then "180" page for Buttress? But then "Chan Kai-man Chan Kai-ning" next.
Lines 392-395: "Chan Kai-man Chan Kai-ning", "**", "161", "28" - "**" maybe separator, pages 161 and 28?
This is extremely messy. The OCR quality is poor.
Given the instruction to "correct unambiguous OCR spelling errors" and "insert ... for missing text", but not to rephrase.
But the primary task: "Reconstruct Sentences and Paragraphs" - but this is tabular data.
I think the best I can do is to present the data as a cleaned-up list in a table, but I must make decisions.
Given the time, I'll produce a Markdown table with three columns (A-Continued, B-Continued, C) by taking the entries in order and distributing them into three columns sequentially (first entry colA, second colB, third colC, fourth colA, etc.). But I need a clean list of entries first.
I'll parse the lines into entries by grouping: a name entry consists of one or more non-page lines until a page line. But page lines are not always clear.
Simpler: Since the original layout is three columns, and the OCR read row-wise, each "row" in the original has three entries. In the OCR output, the entries appear in order: row1col1, row1col2, row1col3, row2col1, row2col2, row2col3, etc. So if I can identify each entry (name+page) in sequence, I can assign them to columns by modulo 3.
But identifying entries is the challenge.
Given the difficulty, perhaps the user expects me to output the text largely as-is but with fixed line breaks and a table? But the instruction says "Return the proofread text in standard Markdown".
Maybe I should output the text as a cleaned-up version of the linear text, with paragraph breaks for each entry? But they said "Use Markdown table syntax to reconstruct tabular data."
I'll attempt to create a table with the three columns by manually parsing the first few rows and then noting the rest is similar? But that would be incomplete.
Given the context, this might be a test. The user provided a sample of OCR output. They want to see how I handle it.
I'll produce a Markdown table with the three columns, but I'll only include the first few rows as a demonstration? But the instruction says to proofread the provided OCR output. The provided OCR output is the entire page. I should process all.
I could write a script in my reasoning to parse, but it's too long.
Alternative: I'll output the proofread text as a single list (two columns: Name, Page) in a table, sorted by the order in the OCR (which is row-major). But that doesn't restore column order.
The instruction: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense (left-to-right, top-to-bottom)." This suggests the OCR read in the wrong order. But here the OCR seems to have read left-to-right, top-to-bottom (row by row). That is one valid reading order. However, for an index, the logical sense is to read down each column. So we should reorder to column-major: all column1 entries, then column2, then column3.
But we don't know where column boundaries are. The headers "A-Continued.", "B-Continued.", "C" indicate three columns. So the entries under each header belong to that column. In the OCR, they are interleaved.
Thus, we need to de-interleave.
Assuming the three columns have equal number of entries (or nearly), we can take the total entries, divide by 3, and assign.
But we need the total entries.
Let me count the entries by identifying page lines.
I'll go through the lines and mark page lines. A page line is a line that consists mainly of digits, commas, dots, maybe "**"? But "**" is not a page.
I'll write a quick algorithm in my head: iterate lines, if line matches /^[\d,\.\s]+$/ then it's a page line. But many name lines have digits (e.g., "B. I.", "L. H. C.", "D.S.O.", "M.A.", "B.SC.", "Edin.", "B.A.", "B.80.", "H. H.", "R. S.", "J. W. M.", "P. W.", "R. P.", "H. J.", "J. D.", "E. F.", "C. H. G.", "F. W.", "C. M.", "A. T.", "S. F.", "C. W.", "V. C.", "J. I.", "J. O.", "N. W.", "A. J. C.", "F. C.", "H. T.", "R. H. J.", "A. O.", "H. R.", "B. M.", "K. F.", "L. H. V.", "N. A.", "G. W.", "J. M.", "B. H.", "W. H.", "F. T.", "J.", "R.", "A.", "C.", "M.", "D.", "E.", "G.", "L.", "N.", "P.", "S.", "T.", "V.", "W.", "Y.", "Z.") but those are initials.
Better: The page numbers are usually at the end of an entry and are numbers like 134, 6, 104, etc. They appear on separate lines in the OCR.
In the line list, many lines are just numbers: 134, 6, 104, 162, 77, 47, 25, 105, 76, 87, 107, 8, 28, 134, 180, 177, 153, 141, 157, 21, 162, 205, 160, 4, 11, 26, 179, 179, 206, 12, 184, 192, 127, 67, 20, 79, 21, 71, 181, 166, 34, 179, 13, 196, 153, 147, 79, 199, 80, 99, 129, 96, 76, 110, 111, 32, 69, 92, 203, 178, 42, 99, 70, 170, 95, 130, 33, 88, 34, 147, 97, 93, 171, 97, 130, 33, 113, 17, 132, 63, 127, 28, 182, 196, 68, 158, 45, 209, 164, 127, 66, 97, 67, 149, 5, 74, 125, 119, 187, 62, 143, 164, 141, 129, 4. 21, 139, 183, 80, 175, 88, 48, 202, 21, 108, 84, 5. 36, 126, 185, 12, 90, 51, 150, 203, 195, 160, 198, 147, 22, 7, 66, 12, 45, 67, 79, 115, 182, 147, 182, 67, 170, 151, 48, 194, 107, 62, 136, 153, 38, 18, 184, 150, 118, 153, 62, 101, 118, 114, 183, 183, 189, 202, 6, 203, 88, 144, 150, 152, 162, 130, 71, 136, 108, 97, 44, 182, 197, 36, 175, 193, 16, 135, 5, 43, 93, 76, 196, 152, 4. 6, 18, 180, 161, 28.
But many of these are not on separate lines in the line list; they are on lines with other text? In my line list, each line is separate. So the page lines are separate.
But some page lines have text like "4. 21", "5. 36", "4. 6", "18 Buttress, E. F." (line 390), "****" (line 393), "A" (line 341), "•" (line 262), "---" (line 311), "·" (line 304), "." (line 183), "-" (line 117), "Në –" (line 272), "...." (line 353), "*****" (line 393).
So the pure numeric lines are pages.
But the entries are interleaved: name lines then page line.
In the line list, the pattern is: some name lines (1 or more), then a page line (numeric), then next name lines, etc.
But because there are three columns interleaved, the sequence is: col1 name(s), col1 page, col2 name(s), col2 page, col3 name(s), col3 page, col1 name(s), col1 page, ...
So if I can segment the line list into entries (name block + page line), then I can assign entries to columns by order.
Let me attempt to segment the line list from line 19 onward (after headers).
I'll go through lines 19-395 and group.
I'll define a page line as a line that matches ^[\d,\.\s]+$ (only digits, commas, periods, spaces). But "4. 21" has a period and space, "5. 36", "4. 6", "6, 203", "110, 111", "143, 164", "160, 198", "4, 11", etc. Also "****" not. "A" not. "•" not. "---" not. "·" not. "." not. "-" not. "Në –" not. "...." not. "18 Buttress, E. F." not.
So I'll consider a line as a page line if it consists only of digits, commas, periods, spaces, and maybe hyphens? But "4. 21" could be "4, 21"? OCR error.
I'll manually segment.
Start at line 19: "Au Sze-bun" (name)
Line 20: "134" (page) -> entry1: name="Au Sze-bun", page="134"
Line 21: "Bickford, B. I." (name)
Line 22: "6" (page) -> entry2: name="Bickford, B. I.", page="6"
Line 23: "Cable, R. E." (name)
Line 24: "104" (page) -> entry3: name="Cable, R. E.", page="104"
Line 25: "Au Tai-yuen" (name)
Line 26: "162" (page) -> entry4
Line 27: "Bidmead, K. A." (name)
Line 28: "77" (page) -> entry5
Line 29: "Cairns, D. G." (name)
Line 30: "47" (page) -> entry6
Line 31: "Au Tse-tsau" (name)
Line 32: "25" (page) -> entry7
Line 33: "Bing-jan, A." (name)
Line 34: "105" (page) -> entry8
Line 35: "Calthrop, L. H. C." (name)
Line 36: "76" (page) -> entry9
Line 37: "Au Wai-ming" (name)
Line 38: "87" (page) -> entry10
Line 39: "Binns, M." (name)
Line 40: "107" (page) -> entry11
Line 41: "Cameroo, M. A." (name)
Line 42: "8" (page) -> entry12
Line 43: "Au Wai-shum" (name)
Line 44: "28" (page) -> entry13
Line 45: "Bird, L. G.," (name part)
Line 46: "D.S.O.," (name part)
Line 47: "134" (page) -> but then line 48: "Colonel" (name part?) This is problematic. The page appears before "Colonel". So maybe the name is "Bird, L. G., D.S.O., Colonel" and page 134. But the OCR placed page in middle. Since the pattern is name then page, but here page appears before last part of name. Could be that "Colonel" is actually the start of next entry? But next is "Campbell, J. G." line 49. "Colonel" might be a title for Bird. So entry14: name="Bird, L. G., D.S.O., Colonel", page="134". But the page line is line 47, and "Colonel" line 48 is after page. So the segmentation fails.
Similarly, line 49: "Campbell, J. G." name, line 50: "180" page -> entry15.
Line 51: "Au Wai-sum" name, line 52: "177" page -> entry16.
Line 53: "Birt, M. D." name, line 54: "153" page -> entry17.
Line 55: "Carr, J. R." name, line 56: "141" page -> entry18.
Line 57: "Au Wochung" name, line 58: "157" page -> entry19.
Line 59: "Bishan Dass" name, line 60: "21" page -> entry20.
Line 61: "Carr, T. W." name, line 62: "162" page -> entry21.
Line 63: "Au Young-chong" name, line 64: "205" page -> entry22.
Line 65: "Bishen Singh" name, line 66: "160" page -> entry23.
Line 67: "Carrie, W. J., M.A.," name part
Line 68: "B.SC., (Edin.)" name part
Line 69: "4, 11" page -> entry24: name="Carrie, W. J., M.A., B.SC., (Edin.)", page="4, 11"
Line 70: "Au Yeung-fan" name, line 71: "26" page -> entry25.
Line 72: "Bishop, C. W. E." name, line 73: "179" page -> entry26.
Line 74: "Carter, E. S." name, line 75: "179" page -> entry27.
Line 76: "Au Yeung-kwong" name, line 77: "206" page -> entry28.
Line 78: "Black, T." name, line 79: "12" page -> entry29.
Line 80: "Casey, E." name, line 81: "184" page -> entry30.
Line 82: "Au Yeung-lam" name, line 83: "192" page -> entry31.
Line 84: "Blake, M." name, line 85: "127" page -> entry32.
Line 86: "Cash, A. I." name, line 87: "67" page -> entry33.
Line 88: "Au Yeung-man" name, line 89: "20" page -> entry34.
Line 90: "Bloor, E." name, line 91: "79" page -> entry35.
Line 92: "Castilho, A. F." name, line 93: "21" page -> entry36.
Line 94: "Awtar Singh" name, line 95: "71" page -> entry37.
Line 96: "Bolt, T." name, line 97: "181" page -> entry38.
Line 98: "Cathie, D. C." name, line 99: "166" page -> entry39.
Line 100: "Ayock, W." name, line 101: "34" page -> entry40.
Line 102: "Bond. G. H." name, line 103: "179" page -> entry41.
Line 104: "Chak Ho-ka" name, line 105: "13" page -> entry42.
Line 106: "Boast, R J." name, line 107: "196" page -> entry43.
Line 108: "Booker, D." name, line 109: "153" page -> entry44.
Line 110: "Chak Ping-ki" name, line 111: "147" page -> entry45.
Line 112: "Booker, F. E. E." name, line 113: "79" page -> entry46.
Line 114: "Chambers, G. J." name, line 115: "199" page -> entry47.
Line 116: "Bachan Singh" name
Line 117: "-" (maybe part of name? or dash)
Line 118: "80" page -> entry48: name="Bachan Singh -", page="80"? Or "Bachan Singh" page 80.
Line 119: "Boota" name, line 120: "99" page -> entry49.
Line 121: "Champelovier, C. T.." name, line 122: "129" page -> entry50.
Line 123: "Bachan Singh I" name, line 124: "96" page -> entry51.
Line 125: "Booth, L. H. V." name, line 126: "76" page -> entry52.
Line 127: "Chan, A." name, line 128: "110, 111" page -> entry53.
Line 129: "Bachan Singh Gill" name, line 130: "32" page -> entry54.
Line 131: "Borges. J. A." name, line 132: "69" page -> entry55.
Line 133: "Chan, B." name, line 134: "92" page -> entry56.
Line 135: "Badan Singh" name, line 136: "203" page -> entry57.
Line 137: "Bottomley, J. H." name, line 138: "178" page -> entry58.
Line 139: "Chan Chak-sze" name, line 140: "42" page -> entry59.
Line 141: "Bagga Singh" name, line 142: "99" page -> entry60.
Line 143: "Bourchier, W. H. C." name, line 144: "70" page -> entry61.
Line 145: "Chan Cheuk" name, line 146: "170" page -> entry62.
Line 147: "Bagh Singh" name, line 148: "95" page -> entry63.
Line 149: "Bowden, G. W." name, line 150: "130" page -> entry64.
Line 151: "Chan Cheuk-sing" name, line 152: "33" page -> entry65.
Line 153: "Bagley. W. J." name, line 154: "88" page -> entry66.
Line 155: "Bowen, J." name, line 156: "34" page -> entry67.
Line 157: "Chan Cheuk-wa, B.A." name, line 158: "147" page -> entry68.
Line 159: "Bahnder Khan I" name, line 160: "97" page -> entry69.
Line 161: "Bowen, N. A." name, line 162: "93" page -> entry70.
Line 163: "Chan Cheung" name, line 164: "171" page -> entry71.
Line 165: "Bahader Khan II" name, line 166: "97" page -> entry72.
Line 167: "Boyd, J. M." name, line 168: "130" page -> entry73.
Line 169: "Chan Cheung-ming..." name, line 170: "33" page -> entry74.
Line 171: "Bailey, B. H." name, line 172: "113" page -> entry75.
Line 173: "Bradley, C. H. G." name, line 174: "17" page -> entry76.
Line 175: "Chan Chi-fat" name, line 176: "132" page -> entry77.
Line 177: "Bailey, W. H." name, line 178: "63" page -> entry78.
Line 179: "Bradley, F. W." name, line 180: "127" page -> entry79.
Line 181: "Chan Chi-hing" name, line 182: "28" page -> entry80.
Line 183: "*." (maybe stray)
Line 184: "Baker, F. T." name, line 185: "182" page -> entry81.
Line 186: "Brailsford, A." name, line 187: "196" page -> entry82.
Line 188: "Chan Chi-ling" name, line 189: "68" page -> entry83.
Line 190: "Baker, J." name, line 191: "158" page -> entry84.
Line 192: "Brailsford, C. M." name, line 193: "45" page -> entry85.
Line 194: "Chan Chi-sun" name, line 195: "209" page -> entry86.
Line 196: "Baker, R." name, line 197: "164" page -> entry87.
Line 198: "Braley, A. T." name, line 199: "127" page -> entry88.
Line 200: "Chan Chieng" name, line 201: "66" page -> entry89.
Line 202: "Bakhtawer Singh Balfour, S. F." name (maybe two names)
Line 203: "97" page -> entry90: name="Bakhtawer Singh Balfour, S. F.", page="97"
Line 204: "Brand, C. W." name, line 205: "67" page -> entry91.
Line 206: "Chan Chik-ting" name, line 207: "149" page -> entry92.
Line 208: "5, 74, 125" page? But no name before. This might be page for previous? But previous already has page. Or it's a page for "Balfour, S. F." if that was separate. But we treated as one name. Could be that "Bakhtawer Singh" and "Balfour, S. F." are two entries: line 202 "Bakhtawer Singh Balfour, S. F." might be two lines merged. But in OCR it's one line. Then line 203 "97" page for Bakhtawer Singh. Then line 204 "Brand, C. W." etc. Then line 208 "5, 74, 125" could be page for Balfour, S. F. But then there is no name line for Balfour. So maybe the name "Balfour, S. F." is missing. We'll treat "5, 74, 125" as a page line without a name? Or it's a stray. I'll assume it's a page for an entry that lost its name. But we have to insert ... for missing name.
Line 209: "Branson, V. C." name, line 210: "119" page -> entry93.
Line 211: "Chan Chin-kun" name, line 212: "187" page -> entry94.
Line 213: "Bumsey, S. F. Barnes, J. I." name (maybe two)
Line 214: "Barnet, J." name
Line 215: "62" page -> entry95: name="Bumsey, S. F. Barnes, J. I. Barnet, J.", page="62"
Line 216: "Brawn, A. O." name, line 217: "143, 164" page -> entry96.
Line 218: "Chan Cho-cheong" name, line 219: "141" page -> entry97.
Line 220: "129" page? stray? Then line 221: "Breen, M. J." name, line 222: "4. 21" page -> entry98.
Line 223: "Chan Chue-kan" name, line 224: "139" page -> entry99.
Line 225: "183" page? stray? Then line 226: "Brennan, J." name, line 227: "80" page -> entry100.
Line 228: "Chan Chuen" name, line 229: "175" page -> entry101.
Line 230: "Barrett, H." name
Line 231: "Barros, J. O." name
Line 232: "Barrow, J." name
Line 233: "88" page -> entry102: name="Barrett, H. Barros, J. O. Barrow, J.", page="88"
Line 234: "Brett, F." name, line 235: "48" page -> entry103.
Line 236: "Chan Chuk-kwan" name, line 237: "202" page -> entry104.
Line 238: "21" page? stray? Then line 239: "Brett, V. N." name, line 240: "108" page -> entry105.
Line 241: "Chan Chun-ip" name, line 242: "84" page -> entry106.
Line 243: "5. 36" page? stray? Then line 244: "Brewer, L." name, line 245: "126" page -> entry107.
Line 246: "Chan Chung-ping" name, line 247: "185" page -> entry108.
Line 248: "Barton, L. A." name
Line 249: "Bascombe, N. W., B.A." name
Line 250: "Bashir Hussain" name
Line 251: "Basto, A. J. C." name
Line 252: "12" page -> entry109: name="Barton, L. A. Bascombe, N. W., B.A. Bashir Hussain Basto, A. J. C.", page="12"
Line 25
116
[ii]
LIST OF OFFICERS.
OFFICE
PAGE
OFFICE
PAGE
OFFICE
PAGE
J
J
J
A-Continued.
B-Continued.
C
Au Sze-bun
134
Bickford, B. I.
6
Cable, R. E.
104
Au Tai-yuen
162
Bidmead, K. A.
77
Cairns, D. G.
47
Au Tse-tsau
25
Bing-jan, A.
105
Calthrop, L. H. C.
76
Au Wai-ming
87
Binns, M.
107
Cameroo, M. A.
8
Au Wai-shum
28
Bird, L. G.,
D.S.O.,
Campbell, J. G.
180
Au Wai-sum
177
134
Colonel
Carr, J. R.
141
Au Wochung
157
Birt, M. D.
153
Carr, T. W.
162
Au Young-chong
205
Bishan Dass
21
Carrie, W. J., M.A.,
Au Yeung-fan
26
Bishen Singh
160
B.SC., (Edin.)
4, 11
Au Yeung-kwong
206
Bishop, C. W. E.
179
Carter, E. S.
179
Au Yeung-lam
192
Black, T.
12
Casey, E.
184
Au Yeung-man
20
Blake, M.
127
Cash, A. I.
67
Awtar Singh
71
Bloor, E.
79
Castilho, A. F.
21
Ayock, W.
34
Boast, R J.
196
Castilho, J. R.
70
Bolt, T.
181
Cathie, D. C.
166
Bond. G. H.
179
Chak Ho-ka
13
B
Booker, D.
153
Chak Ping-ki
147
Booker, F. E. E.
79
Chambers, G. J.
199
Bachan Singh
-
80
Boota
99
Champelovier, C. T..
129
Bachan Singh I
96
Booth, L. H. V.
76
Chan, A.
110, 111
Bachan Singh Gill
32
Borges. J. A.
69
Chan, B.
92
Badan Singh
203
Bottomley, J. H.
178
Chan Chak-sze
42
Bagga Singh
99
Bourchier, W. H. C.
70
Chan Cheuk
170
Bagh Singh
95
Bowden, G. W.
130
Chan Cheuk-sing
33
Bagley. W. J.
88
Bowen, J.
34
Chan Cheuk-wa, B.A.
147
Bahnder Khan I
97
Bowen, N. A.
93
Chan Cheung
171
Bahader Khan II
97
Boyd, J. M.
130
Chan Cheung-ming...
33
Bailey, B. H.
113
Bradley, C. H. G.
17 Chan Chi-fat
132
Bailey, W. H.
63
Bradley, F. W.
127
Chan Chi-hing
28
*.
Baker, F. T.
182
Brailsford, A.
196
Chan Chi-ling
68
Baker, J.
158
Brailsford, C. M.
45
Chan Chi-sun
209
Baker, R.
164
Braley, A. T.
127
Chan Chieng
66
Bakhtawer Singh Balfour, S. F.
97
Brand, C. W.
67
Chan Chik-ting
149
5, 74, 125
Branson, V. C.
119
Chan Chin-kun
187
Bumsey, S. F. Barnes, J. I.
Barnet, J.
62
Brawn, A. O.
143, 164
Chan Cho-cheong
141
129
Breen, M. J.
Chan Chue-kan
139
183
Brennan, J.
80
Chan Chuen
175
Barrett, H.
Barros, J. O.
Barrow, J.
88
Brett, F.
48
Chan Chuk-kwan
202
21
Brett, V. N.
108
Chan Chun-ip
84
Brewer, L.
126
Chan Chung-ping
185
Barton, L. A.
Bascombe, N. W., B.A.
Bashir Hussain
Basto, A. J. C.
12
Brimblecombe, F. C.
90
Chan Fo-po
51
150
Broadbridge, N.
203
Chan Fook-chor
195
•
160, 198
Broadbridge, S. A.
147
Chan Fuk-chi
22
7
Brooks, H. T.
66
Chan Fuk-him
Në –
12
Bates, R. A.
45
Brooks, R. H. J.
67
Chan Fung-cheung
79
Bau Tau Zung
115
Brooksbank, A.
182
Chau Fung-kee
147
Beach. J. S.
182
Brown, E.
67
Chan Ha
170
Beattie. A. F.
151
Brown, E. F.
48
Chan Hang
194
Beattie, C.
107
Brown, H. C.
62
Chan Hi
136
·
Beavis, E. M.
153
Brown, J. W. M.
38
Chan Hi-wo
18
---
Bebbington, N. J.
184
Brown, P. W., B.A....
150
Chan Hing-man
118
Boddow, H. H., B.80.
153
Brown, R. P.
62
Chan Hing-tong
101
Begbie, R. S.
118
Brown, S.
114
Chan Hiü-fung
183
Bell, R. S.
183
Bruce, W.
189
Chan Hok
202
Bendall. J. A.
6, 203
Buchanan, G. L.
88
Chau Hon-shan
A
144
Benington, K. F.
150
Buckwell, J. W.
152
Chan Huen-pak
162
Benwell, V. M.
130
Budden. B.
71
Chan Hung ....
136
Berkley, F.
108
Buller, W. J.
97
Chan Hung-tat
44
Best, H. J.
182
Bunn, E.
197
Chan Im-larn
36
Bevan, R. S.
175
Burford, A. O.
193
Chan Tu-fong
16
Bhagat Singh
135
Burgess. C. B.
5
Chau Iu-kwong
43
Bhan Singh
93
Burlingham, D.
76
Chan Ka-leung
196
Bicheno, B. M.
152
Butters, H. R.
Bickerstaff, J. D.
18 Buttress, E. F.
180
Chan Kai-man Chan Kai-ning
*****
161
28
No comments yet.
Private notes are available after approval.