The user wants me to proofread OCR output of a historical Hong Kong government document (likely a civil service list or establishment register). The text appears to be a table of Education Department staff with columns: Date of Appointment, Office, Name, Date of First Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933.
The OCR has broken the table structure. I need to reconstruct it as a Markdown table. The data seems to span multiple lines per row. Let me parse carefully.
First, identify the headers from the first lines:
"Date of Appointment. OFFICE, NAME. of Appointment. Authority. Annual Salary. EDUCATION DEPARTMENT,—Continued. House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose. Leave taken during 1933."
Then rows:
Actually the table seems to have two date columns: "Date of Appointment" (current appointment?) and "Date of First Appointment". The OCR shows at the bottom a list of dates under "Date of First Appointment." and "ART 274 (J 144)" then a list of dates: 25th February, 1929. 1st September, 1915. 1st March, 1920. 1st June, 1915. 17th February, 1919. Do. 1st March. 1923. Do. 2nd April, 1928. 10th April, 1903. 20th June, 1922. 7th January, 1924. 23rd March, 1920. 24th April, 1923. 3rd April, 1923. 1st September, 1926.
These correspond to each row's "Date of First Appointment". The rows in order:
That's 16 rows. The date list has 16 entries? Let's count:
Yes 16 entries.
Now the other columns: "Date of Appointment" (current appointment date) appears in the main body:
"Inspector of Vernacular Schools, Liang Sai-wa, B.A.
Sub-Inspector of Vernacular
Lau Shuk-chong.
Schools, Urban District,
Do.,
Do..
Sub-Inspector of Vernacular
Lam Tun-wing.
25th February, 1929. 1st January, 1927.
Do.
C.S.O. 5006 of 1929.
$5,400 $25 from the Technical Institute.
C.S.O. 888 of 1926.
3,800
Do.
3,200
Chan Tni-shang,
Do.
Do.
3,200
Lai Kuk-yoen.
17th February,
C.S.O. 6502 of 1910,
2,550
Schools for New Territories,
Do.,
Hui Tat-cheung.
Do.
2,400
Do.,
Chan Kam-tong.
Do.,
Do..
Chan Yuk-cheung.
To Chung-pok.
1st January, 1927. Do.
Do.
1,950
Do.
1,950
2nd April,
1928.
C.S.O. 4841 of 1928.
1,650
Class I Clerk,
(J. C. S.)
S. M. Suddick.
1st January, 1929.
C.S.O. 5254 of 1929.
4,300
Class IV Clerk,
( do.)
Leung Tak-wa,
Do.
( do. )
Ho Ka-chun.
(1)
Class V Clerk,
( do.)
Chan Hon-shan.
1st January, 1930. 1st January, 1933. 1st May, 1926.
C.S.O. 5254 of 1930,
1,800 $60 Shorthand Allowance,
C.S.O. 5254 of 1933,
1,500
C.S.O. 3646 of 1926.
1,400
I
Do,
( do.)
Luk Hok-king.
1st January, 1929.
******
1,400
Do.,
( do. )
To Man-to.
1st January, 1931.
C.S.O. 5255 of 1931.
1,300
Do.,
( do.)
Chan Sik-pui,
1st December, 1931.
C.S.O. 5255 of 1931.
1,300"
This is messy. Let's reconstruct row by row.
I'll create a Markdown table with columns: Date of Appointment, Office, Name, Date of First Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933.
But the OCR doesn't show "Leave taken during 1933" column data; maybe it's empty or not captured. The header includes it but no data visible. Might be blank. I'll include column but leave empty.
Also "House or Quarters, and Allowances..." column: only for Liang Sai-wa: "$25 from the Technical Institute." and for Chan Hon-shan: "$60 Shorthand Allowance". Others maybe none.
Now parse each row:
Row 1: Inspector of Vernacular Schools, Liang Sai-wa, B.A.
Row 2: Sub-Inspector of Vernacular Schools, Urban District, Lau Shuk-chong.
Row 3: Sub-Inspector of Vernacular Schools, Urban District, Lam Tun-wing.
Let's read sequentially:
After "Lau Shuk-chong." line, next line "Schools, Urban District," then "Do.," then "Do.." then "Sub-Inspector of Vernacular" then "Lam Tun-wing." then "25th February, 1929. 1st January, 1927." then "Do." then "C.S.O. 5006 of 1929." then "$5,400 $25 from the Technical Institute." then "C.S.O. 888 of 1926." then "3,800" then "Do." then "3,200" then "Chan Tni-shang," then "Do." then "Do." then "3,200" then "Lai Kuk-yoen." then "17th February," then "C.S.O. 6502 of 1910," then "2,550" then "Schools for New Territories," then "Do.," then "Hui Tat-cheung." then "1927. Do." then "Do." then "2,400" then "Do.," then "Chan Kam-tong." then "Do.," then "Do.." then "Chan Yuk-cheung." then "To Chung-pok." then "1st January, 1927. Do." then "Do." then "1,950" then "Do." then "1,950" then "2nd April," then "1928." then "C.S.O. 4841 of 1928." then "1,650" then "Class I Clerk," then "(J. C. S.)" then "S. M. Suddick." then "1st January, 1929." then "C.S.O. 5254 of 1929." then "4,300" then "Class IV Clerk," then "( do.)" then "Leung Tak-wa," then "Do." then "( do. )" then "Ho Ka-chun." then "(1)" then "Class V Clerk," then "( do.)" then "Chan Hon-shan." then "1st January, 1930. 1st January, 1933. 1st May, 1926." then "C.S.O. 5254 of 1930," then "1,800 $60 Shorthand Allowance," then "C.S.O. 5254 of 1933," then "1,500" then "C.S.O. 3646 of 1926." then "1,400" then "I" then "Do," then "( do.)" then "Luk Hok-king." then "1st January, 1929." then "******" then "1,400" then "Do.," then "( do. )" then "To Man-to." then "1st January, 1931." then "C.S.O. 5255 of 1931." then "1,300" then "Do.," then "( do.)" then "Chan Sik-pui," then "1st December, 1931." then "C.S.O. 5255 of 1931." then "1,300"
This is a linear OCR reading of a table, likely row by row but with line breaks within cells. The pattern: each row has: Office, Name, Date of Appointment, Authority, Annual Salary, Allowances, (maybe Date of First Appointment later). But the Date of First Appointment appears at the end as a separate list.
Better approach: The original document is a "Civil Service List" or "Establishment Register" with columns: Date of Appointment, Office, Name, Date of First Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933.
The OCR has read the page in a columnar fashion but messed up. The bottom list "Date of First Appointment." with "ART 274 (J 144)" and then dates is likely the "Date of First Appointment" column for each row.
Thus we have 16 rows. Let's assign each row an index and fill columns from the linear text.
I'll write a script mentally but let's do manually.
First, identify each row's office and name:
Now extract Date of Appointment (current appointment) for each from the linear text. The linear text after each name seems to have a date. Let's locate:
After "Liang Sai-wa, B.A." the next date is "25th February, 1929." (but that appears later after Lam Tun-wing? Actually the text: "25th February, 1929. 1st January, 1927." appears after Lam Tun-wing. Hmm.
Let's segment the linear text by rows. The OCR likely read the table row by row, but each row's cells are spread across lines. The pattern: For each row, the cells are: Office, Name, Date of Appointment, Authority, Annual Salary, Allowances, (maybe Date of First Appointment). But the Date of First Appointment is separate at bottom.
Look at the start: "Inspector of Vernacular Schools, Liang Sai-wa, B.A." then "Sub-Inspector of Vernacular" then "Lau Shuk-chong." then "Schools, Urban District," then "Do.," then "Do.." then "Sub-Inspector of Vernacular" then "Lam Tun-wing." then "25th February, 1929. 1st January, 1927." then "Do." then "C.S.O. 5006 of 1929." then "$5,400 $25 from the Technical Institute." then "C.S.O. 888 of 1926." then "3,800" then "Do." then "3,200" then "Chan Tni-shang," then "Do." then "Do." then "3,200" then "Lai Kuk-yoen." then "17th February," then "C.S.O. 6502 of 1910," then "2,550" then "Schools for New Territories," then "Do.," then "Hui Tat-cheung." then "1927. Do." then "Do." then "2,400" then "Do.," then "Chan Kam-tong." then "Do.," then "Do.." then "Chan Yuk-cheung." then "To Chung-pok." then "1st January, 1927. Do." then "Do." then "1,950" then "Do." then "1,950" then "2nd April," then "1928." then "C.S.O. 4841 of 1928." then "1,650" then "Class I Clerk," then "(J. C. S.)" then "S. M. Suddick." then "1st January, 1929." then "C.S.O. 5254 of 1929." then "4,300" then "Class IV Clerk," then "( do.)" then "Leung Tak-wa," then "Do." then "( do. )" then "Ho Ka-chun." then "(1)" then "Class V Clerk," then "( do.)" then "Chan Hon-shan." then "1st January, 1930. 1st January, 1933. 1st May, 1926." then "C.S.O. 5254 of 1930," then "1,800 $60 Shorthand Allowance," then "C.S.O. 5254 of 1933," then "1,500" then "C.S.O. 3646 of 1926." then "1,400" then "I" then "Do," then "( do.)" then "Luk Hok-king." then "1st January, 1929." then "******" then "1,400" then "Do.," then "( do. )" then "To Man-to." then "1st January, 1931." then "C.S.O. 5255 of 1931." then "1,300" then "Do.," then "( do.)" then "Chan Sik-pui," then "1st December, 1931." then "C.S.O. 5255 of 1931." then "1,300"
This is confusing. Perhaps the OCR read the table column by column? But the presence of "Do." suggests ditto marks.
Let's think of the original table layout. Typically such lists have columns:
The OCR output seems to have the Office and Name combined in first column? Actually the header: "Date of Appointment. OFFICE, NAME. of Appointment. Authority. Annual Salary. EDUCATION DEPARTMENT,—Continued. House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose. Leave taken during 1933."
So the columns are: Date of Appointment, Office, Name, Date of First Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave.
But the OCR has "Date of Appointment." then "OFFICE, NAME." then "of Appointment." (maybe "Date of First Appointment"?), then "Authority.", "Annual Salary.", then "EDUCATION DEPARTMENT,—Continued." (that's a section header), then "House or Quarters...", "Leave taken during 1933."
Thus the table has 8 columns.
Now the data rows: each row should have 8 values.
The linear OCR text after header seems to list values row by row but with line breaks. Let's try to parse by matching known patterns.
We have 16 rows. The bottom list gives Date of First Appointment for each row (16 dates). Good.
Now we need to extract for each row: Date of Appointment (col1), Office (col2), Name (col3), Authority (col5), Annual Salary (col6), Allowances (col7), Leave (col8). Date of First Appointment (col4) from bottom list.
Let's go through linear text and assign to rows sequentially.
I'll create an array of rows, each with fields.
Start after header. The first data likely: "Inspector of Vernacular Schools, Liang Sai-wa, B.A." That's Office and Name for row1.
Next: "Sub-Inspector of Vernacular" then "Lau Shuk-chong." then "Schools, Urban District," then "Do.," then "Do.." This seems like row2: Office: "Sub-Inspector of Vernacular Schools, Urban District", Name: "Lau Shuk-chong". The "Do., Do.." might be ditto for Office? Actually "Do." means same as above. But the office for row2 is same as row1? No, row1 is Inspector, row2 is Sub-Inspector. So "Do." might be for something else.
Maybe the OCR read the table vertically? Let's look at the structure of the original image: It might be a two-page spread with columns. But we have to do our best.
Given the difficulty, perhaps the best is to reconstruct the table as a Markdown table with the data we can infer, using the bottom dates for Date of First Appointment, and the linear text for other columns by aligning with the row order.
We have 16 rows. Let's list the linear text tokens that look like data for each row in order.
I'll write the linear text as a sequence of lines (as given). But the user gave a block of text. I'll split by lines as they appear in the prompt.
The prompt text lines:
"Date
OFFICE,
NAME.
of Appointment.
Authority.
Auunal Salary.
EDUCATION DEPARTMENT,—Continued.
House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose.
Leave taken during 1933.
Inspector of Vernacular Schools, Liang Sai-wa, B.A.
Sub-Inspector of Vernacular
Lau Shuk-chong.
Schools, Urban District,
Do.,
Do..
Sub-Inspector of Vernacular
Lam Tun-wing.
25th February, 1929. 1st January, 1927.
Do.
C.S.O. 5006 of 1929.
$5,400 $25 from the Technical Institute.
C.S.O. 888 of 1926.
3,800
Do.
3,200
Chan Tni-shang,
Do.
Do.
3,200
Lai Kuk-yoen.
17th February,
C.S.O. 6502 of 1910,
2,550
Schools for New Territories,
Do.,
Hui Tat-cheung.
Do.
2,400
Do.,
Chan Kam-tong.
Do.,
Do..
Chan Yuk-cheung.
To Chung-pok.
1st January, 1927. Do.
Do.
1,950
Do.
1,950
2nd April,
1928.
C.S.O. 4841 of 1928.
1,650
Class I Clerk,
(J. C. S.)
S. M. Suddick.
1st January, 1929.
C.S.O. 5254 of 1929.
4,300
Class IV Clerk,
( do.)
Leung Tak-wa,
Do.
( do. )
Ho Ka-chun.
(1)
Class V Clerk,
( do.)
Chan Hon-shan.
1st January, 1930. 1st January, 1933. 1st May, 1926.
C.S.O. 5254 of 1930,
1,800 $60 Shorthand Allowance,
C.S.O. 5254 of 1933,
1,500
C.S.O. 3646 of 1926.
1,400
I
Do,
( do.)
Luk Hok-king.
1st January, 1929.
******
1,400
Do.,
( do. )
To Man-to.
1st January, 1931.
C.S.O. 5255 of 1931.
1,300
Do.,
( do.)
Chan Sik-pui,
1st December, 1931.
C.S.O. 5255 of 1931.
1,300
****
Date of First Appointment.
ART
274
(J 144 )
25th February,
1st March. 1923.
Do.
2nd April, 1928. 10th April, 1903. 20th June, 1922.
7th January,
1923.
1st September, 1926."
Now, the header lines are separate. Then data lines. The "*" and "***" are likely separators or OCR artifacts.
The "Date of First Appointment." section at the end lists 16 dates.
Let's number the rows 1-16 as per the order of names appearing.
Names in order of appearance:
That's 16 names.
Now, for each, we need to extract:
We have the bottom list dates in same order? Likely yes. The bottom list:
That's 16 entries.
Now assign to rows 1-16.
Row1: Date of First Appointment = 25th February, 1929
Row2: 1st September, 1915
Row3: 1st March, 1920
Row4: 1st June, 1915
Row5: 17th February, 1919
Row6: 17th February, 1919 (Do.)
Row7: 1st March, 1923
Row8: 1st March, 1923 (Do.)
Row9: 2nd April, 1928
Row10: 10th April, 1903
Row11: 20th June, 1922
Row12: 7th January, 1924
Row13: 23rd March, 1920
Row14: 24th April, 1923
Row15: 3rd April, 1923
Row16: 1st September, 1926
Now, for each row, find Date of Appointment, Authority, Salary, Allowances from the linear text.
The linear text after the names seems to interleave these fields. Let's try to parse the linear text as a sequence of fields for each row in order.
After the header, the first data line: "Inspector of Vernacular Schools, Liang Sai-wa, B.A." That's row1 Office+Name.
Next lines: "Sub-Inspector of Vernacular" "Lau Shuk-chong." "Schools, Urban District," "Do.," "Do.." This seems like row2 Office+Name but split. Actually "Sub-Inspector of Vernacular Schools, Urban District" is the office, "Lau Shuk-chong" name. The "Do., Do.." might be ditto for something else? Maybe the OCR captured the "Date of Appointment" column for row1 and row2? Let's see.
After "Do.." we have "Sub-Inspector of Vernacular" "Lam Tun-wing." That's row3.
Then "25th February, 1929. 1st January, 1927." These are likely Date of Appointment for row1 and row2? Row1 Date of Appointment = 25th February, 1929. Row2 Date of Appointment = 1st January, 1927.
Then "Do." - maybe Date of Appointment for row3? "Do." means same as previous? But previous is 1st January, 1927. So row3 Date of Appointment = 1st January, 1927.
Then "C.S.O. 5006 of 1929." - Authority for row1.
Then "$5,400 $25 from the Technical Institute." - Salary and Allowances for row1.
Then "C.S.O. 888 of 1926." - Authority for row2.
Then "3,800" - Salary for row2.
Then "Do." - Allowances for row2? "Do." might mean no allowances? Or same as previous? But previous allowance was "$25 from the Technical Institute." That seems specific to row1. "Do." might indicate "Nil" or "None". But in these lists, "Do." often means "Ditto" for the column above. However, the allowance column might be empty for most. The "Do." could be for the Allowances column for row2, meaning same as row1? But that seems unlikely. Maybe "Do." is for the Authority column? But we already have Authority for row2. Let's see the pattern.
After "3,800" we have "Do." then "3,200". That "Do." might be Authority for row3? But row3's authority might be same as row2? "C.S.O. 888 of 1926." So "Do." could be Authority for row3. Then "3,200" is Salary for row3.
Then "Chan Tni-shang," - Name for row4.
Then "Do." - Office for row4? "Do." means same office as previous? Previous office was "Sub-Inspector of Vernacular Schools, Urban District". So row4 same office.
Then "Do." - Date of Appointment for row4? Same as previous? Previous Date of Appointment for row3 was 1st January, 1927 (from "Do."). So row4 Date of Appointment = 1st January, 1927.
Then "3,200" - Salary for row4.
Then "Lai Kuk-yoen." - Name for row5.
Then "17th February," - Date of Appointment for row5? "17th February," likely 1927? But the line "17th February," then next line "C.S.O. 6502 of 1910," then "2,550". So row5: Date of Appointment = 17th February, 1927? But the year not given. Could be 1927. The bottom list shows Date of First Appointment for row5 = 17th February, 1919. So Date of Appointment (current) might be 17th February, 1927? The OCR shows "17th February," then next line "C.S.O. 6502 of 1910," which is Authority. Then "2,550" Salary.
Then "Schools for New Territories," - Office for row5? Actually row5 is "Sub-Inspector of Vernacular Schools, New Territories". The office appears after the name? The order might be: Name, Date of Appointment, Authority, Salary, Office? That seems odd.
Let's look at the header order: Date of Appointment, Office, Name, Date of First Appointment, Authority, Annual Salary, Allowances, Leave.
But the linear OCR might not follow that order because of column reading.
Given the complexity, perhaps the best is to present the data in a Markdown table with the columns as per header, and fill in what we can deduce, leaving blank where uncertain. But the user expects a proofread text, not necessarily a perfect reconstruction. However, the instruction: "Reconstruct the table using Markdown table syntax." So we must produce a table.
Given the OCR is messy, we can try to reconstruct the table as it would appear in the original document. Since this is a known historical document (Hong Kong Civil Service List 1933?), we might know the exact format. But we must work only from the OCR.
Alternative: The OCR text might be from a PDF where the table is split across pages. The "EDUCATION DEPARTMENT,—Continued." suggests continuation.
Maybe we can treat the linear text as a series of rows where each row's data appears in a block. Let's try to group by the appearance of "C.S.O." which indicates Authority.
List all "C.S.O." entries in order:
That's 10 authorities, but we have 16 rows. Some rows share authority (ditters). The "Do." likely indicates same authority as previous.
Similarly, salaries:
$5,400
3,800
3,200
3,200
2,550
2,400
1,950
1,950
1,650
4,300
1,800
1,500
1,400
1,400
1,300
1,300
That's 16 salaries. Good.
Allowances:
$25 from the Technical Institute. (for row1)
$60 Shorthand Allowance (for row13? Actually appears with 1,800 salary for Chan Hon-shan)
Others maybe none.
Date of Appointment (current) from linear text: we have dates:
25th February, 1929
1st January, 1927
Do. (1st January, 1927)
17th February, (1927?)
1st January, 1927. Do.
2nd April, 1928
1st January, 1929
Do. (for Leung Tak-wa? maybe 1st January, 1929?)
1st January, 1930. 1st January, 1933. 1st May, 1926. (three dates for Chan Hon-shan? Possibly Date of Appointment, Date of First Appointment? But Date of First Appointment is separate.)
1st January, 1929 (for Luk Hok-king)
1st January, 1931
1st December, 1931
Let's list them in order as they appear in linear text after the names:
After row3 (Lam Tun-wing) we see "25th February, 1929. 1st January, 1927." Then "Do." Then later "17th February," then "1927. Do." then "1st January, 1927. Do." then "2nd April, 1928." then "1st January, 1929." then "Do." then "1st January, 1930. 1st January, 1933. 1st May, 1926." then "1st January, 1929." then "1st January, 1931." then "1st December, 1931."
That's 13 date entries? But we have 16 rows. Some rows might have "Do." for Date of Appointment.
Let's align with rows:
Row1: Date of Appointment = 25th February, 1929
Row2: Date of Appointment = 1st January, 1927
Row3: Date of Appointment = Do. (1st January, 1927)
Row4: Date of Appointment = ? Not explicitly, but after row3 we have "Chan Tni-shang, Do. Do. 3,200". The two "Do." might be Office and Date of Appointment? Actually "Chan Tni-shang," then "Do." (Office), "Do." (Date of Appointment), "3,200" (Salary). So Row4 Date of Appointment = Do. (1st January, 1927)
Row5: "Lai Kuk-yoen. 17th February, C.S.O. 6502 of 1910, 2,550" So Row5 Date of Appointment = 17th February, 1927? (year missing, but likely 1927)
Row6: "Schools for New Territories, Do., Hui Tat-cheung. 1927. Do. Do. 2,400" This is messy. "Schools for New Territories" might be Office for row5? Actually row5 office is "Sub-Inspector of Vernacular Schools, New Territories". The text "Schools for New Territories," appears after row5's salary? Let's see: after "2,550" we have "Schools for New Territories," then "Do.," then "Hui Tat-cheung." then "1927. Do." then "Do." then "2,400". So row6: Name = Hui Tat-cheung, Office = Sub-Inspector of Vernacular Schools, New Territories (from "Schools for New Territories," and "Do."), Date of Appointment = 1927? "1927. Do." maybe Date of Appointment = 1st January, 1927? "Do." again for something.
Row7: "Do., Chan Kam-tong. Do., Do.. Chan Yuk-cheung." This suggests row7: Chan Kam-tong, row8: Chan Yuk-cheung. Offices same (Do.). Dates? "Do., Do.." maybe Date of Appointment = Do. (1st January, 1927). Salaries: later "1,950 Do. 1,950" for row7 and row8.
Row9: "To Chung-pok. 1st January, 1927. Do. Do. 1,950 Do. 1,950" Wait that seems duplicate. Actually after Chan Yuk-cheung, we have "To Chung-pok. 1st January, 1927. Do. Do. 1,950 Do. 1,950". That might be row9: To Chung-pok, Date of Appointment = 1st January, 1927, Salary = 1,950. But we already have two 1,950 for row7 and row8. The salaries list shows 1,950 appears twice (for row7 and row8). Row9 salary is 1,650? Actually from salary list: after 2,400 (row6), we have 1,950, 1,950, 1,650. So row7: 1,950, row8: 1,950, row9: 1,650. But the text shows "1,950 Do. 1,950" after To Chung-pok? Let's read: "To Chung-pok. 1st January, 1927. Do. Do. 1,950 Do. 1,950". That might be for row9 and row10? But row10 is Class I Clerk with salary 4,300.
I'm getting lost. Given the time, perhaps the best is to produce a Markdown table with the columns and fill in the data as per the typical structure, using the information we have, and note uncertainties with ... but the instruction says: "if a word or phrase is clearly missing due to OCR damage, insert ... in its place." But we are not to add commentary.
Given the difficulty, I'll attempt to reconstruct the table as faithfully as possible by aligning the 16 rows with the 16 salaries, 16 first appointment dates, and the other fields from the linear text in order.
Let's create a structured parsing by writing a small mental script. But as a language model, I can simulate.
I'll assign each row an index and extract fields from the linear text by scanning sequentially and matching patterns.
Define a list of rows in order of names. For each row, we need: DateApp, Office, Name, DateFirstApp, Authority, Salary, Allowances, Leave.
We have Names in order. We have DateFirstApp from bottom list in same order.
Now, the linear text after header contains the other fields in some order. Let's tokenize the linear text (excluding header and bottom list) into a list of tokens (lines). Then we can try to map tokens to rows.
But the linear text includes the names again. So we can use the names as anchors.
Let's write the linear text as a single string and split by newline? The user provided with line breaks. I'll use the lines as given.
Lines after header (starting from "Inspector of Vernacular Schools, Liang Sai-wa, B.A."):
Now, we have 16 names. Let's locate each name in these lines:
Row1: Liang Sai-wa, B.A. - line1
Row2: Lau Shuk-chong - line3
Row3: Lam Tun-wing - line8
Row4: Chan Tni-shang - line17
Row5: Lai Kuk-yoen - line21
Row6: Hui Tat-cheung - line27
Row7: Chan Kam-tong - line32
Row8: Chan Yuk-cheung - line35
Row9: To Chung-pok - line36
Row10: S. M. Suddick - line48
Row11: Leung Tak-wa - line54
Row12: Ho Ka-chun - line57
Row13: Chan Hon-shan - line61
Row14: Luk Hok-king - line72
Row15: To Man-to - line78
Row16: Chan Sik-pui - line84
Good.
Now, for each row, the fields appear around these lines. The table columns are fixed. The original table likely has each row spanning multiple lines in the OCR because of column width. But we can assume that the data for each row appears in the same relative order in the linear text.
Given the header order: Date of Appointment, Office, Name, Date of First Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave.
But the linear text seems to have Office and Name together in some lines, then later Date of Appointment, Authority, Salary, Allowances.
Let's try to see the pattern for row1 (Liang Sai-wa). After line1, the next lines are for row2 and row3 names. Then line9: "25th February, 1929. 1st January, 1927." This line likely contains Date of Appointment for row1 and row2. Since row1 is first, its Date of Appointment = 25th February, 1929. Row2 = 1st January, 1927.
Line10: "Do." - likely Date of Appointment for row3 = same as row2 (1st January, 1927).
Line11: "C.S.O. 5006 of 1929." - Authority for row1.
Line12: "$5,400 $25 from the Technical Institute." - Salary and Allowances for row1.
Line13: "C.S.O. 888 of 1926." - Authority for row2.
Line14: "3,800" - Salary for row2.
Line15: "Do." - Allowances for row2? Could be "Do." meaning no allowances? Or same as row1? But row1 had allowance. In these lists, "Do." in allowance column often means "Nil" or "None". But we'll interpret as "Do." (ditto) but that might be wrong. However, the allowance column for row2 might be empty. The "Do." might be for the Allowances column for row2, indicating "Do." (i.e., same as above? But above is "$25 from the Technical Institute." That seems specific. Usually "Do." in allowance column means "No allowance". But we'll keep as "Do.".
Line16: "3,200" - Salary for row3.
But wait, row3's authority? Not yet seen. After line16, line17 is "Chan Tni-shang," which is row4 name. So row3's authority might be "Do." from line15? But line15 is "Do." after row2's salary. Could be authority for row3? Let's see: The pattern might be: For each row: Authority, Salary, Allowances. But row1 had Authority, Salary+Allowances. Row2 had Authority, Salary, Allowances (Do.). Row3 might have Authority (Do.), Salary (3,200), Allowances (maybe next Do.). But after 3,200 we have row4 name. So row3's allowance missing? Maybe row3's allowance is "Do." from line15? But line15 is before row3's salary. Hmm.
Let's look at row4: line17 "Chan Tni-shang," line18 "Do." line19 "Do." line20 "3,200". So for row4: Office? "Do." (same as row3 office), Date of Appointment? "Do." (same as row3), Salary 3,200. Authority and Allowances not shown? Maybe they are "Do." from previous? But row4's authority might be same as row3? Row3's authority not explicit.
Row5: line21 "Lai Kuk-yoen." line22 "17th February," line23 "C.S.O. 6502 of 1910," line24 "2,550". So row5: Date of Appointment = 17th February, 1927? (year missing), Authority = C.S.O. 6502 of 1910, Salary = 2,550. Allowances? Not shown.
Then line25 "Schools for New Territories," line26 "Do.," line27 "Hui Tat-cheung." line28 "1927. Do." line29 "Do." line30 "2,400". This seems to be row6: Office = "Schools for New Territories" (i.e., Sub-Inspector of Vernacular Schools, New Territories), Name = Hui Tat-cheung, Date of Appointment = 1927? "1927. Do." maybe Date of Appointment = 1st January, 1927? "Do." again, Salary = 2,400. Authority? Not shown, maybe "Do." from previous? But previous authority was C.S.O. 6502 of 1910. Could be same.
Row7: line31 "Do.," line32 "Chan Kam-tong." line33 "Do.," line34 "Do.." line35 "Chan Yuk-cheung." So row7 and row8 names. Offices "Do." (same as row6). Dates "Do." "Do.." maybe Date of Appointment = same as row6 (1st January, 1927). Salaries later.
Row9: line36 "To Chung-pok." line37 "1st January, 1927. Do." line38 "Do." line39 "1,950" line40 "Do." line41 "1,950". This is confusing: two salaries 1,950. Row9 and row10? But row10 is Suddick with salary 4,300. So maybe row9: To Chung-pok, Date of Appointment = 1st January, 1927, Salary = 1,950. Row10? But row10 is Suddick. The second 1,950 might be for row10? But row10 salary is 4,300. So perhaps the two 1,950 are for row7 and row8 (Chan Kam-tong and Chan Yuk-cheung). And row9 To Chung-pok has salary 1,650 (line45). Let's check salary list: we have salaries: 5400, 3800, 3200, 3200, 2550, 2400, 1950, 1950, 1650, 4300, 1800, 1500, 1400, 1400, 1300, 1300. That's 16 salaries. The order of salaries in linear text: line12: 5400, line14: 3800, line16: 3200, line20: 3200, line24: 2550, line30: 2400, line39: 1950, line41: 1950, line45: 1650, line51: 4300, line64: 1800, line66: 1500, line68: 1400, line75: 1400, line81: 1300, line87: 1300. That matches 16 salaries in order of rows 1-16. Good! So the salaries appear in the linear text in row order. Let's verify the row order matches the name order.
Row1 salary at line12 (5400)
Row2 salary at line14 (3800)
Row3 salary at line16 (3200)
Row4 salary at line20 (3200)
Row5 salary at line24 (2550)
Row6 salary at line30 (2400)
Row7 salary at line39 (1950)
Row8 salary at line41 (1950)
Row9 salary at line45 (1650)
Row10 salary at line51 (4300)
Row11 salary at line64 (1800)
Row12 salary at line66 (1500)
Row13 salary at line68 (1400)
Row14 salary at line75 (1400)
Row15 salary at line81 (1300)
Row16 salary at line87 (1300)
Perfect. So the linear text presents salaries in row order. Good.
Now, the other fields: Date of Appointment, Authority, Allowances also appear in row order? Let's see.
Date of Appointment: line9 has two dates: "25th February, 1929. 1st January, 1927." That's for row1 and row2. Line10: "Do." for row3. Then for row4? Line18 and 19 are "Do." "Do." after row4 name. Those might be Date of Appointment and something else. But row4's Date of Appointment might be "Do." (same as row3). Row5: line22 "17th February," (date). Row6: line28 "1927. Do." (maybe date). Row7: line33 "Do." (date). Row8: line34 "Do.." (date). Row9: line37 "1st January, 1927. Do." (two dates? maybe row9 and row10? But row10 date is line49 "1st January, 1929."). Row10: line49 "1st January, 1929." Row11: line55 "Do." (date). Row12: line58 "(1)"? Not a date. Row13: line62 "1st January, 1930. 1st January, 1933. 1st May, 1926." three dates. Row14: line73 "1st January, 1929." Row15: line79 "1st January, 1931." Row16: line85 "1st December, 1931."
But we have 16 rows. Let's list Date of Appointment for each row from the linear text in order of salaries (which match row order). We'll need to extract the date that corresponds to each row.
The dates appear interspersed. Let's find all date-like lines in the linear text in order:
Line9: "25th February, 1929. 1st January, 1927." (two dates)
Line10: "Do."
Line22: "17th February,"
Line28: "1927. Do."
Line33: "Do."
Line34: "Do.."
Line37: "1st January, 1927. Do."
Line49: "1st January, 1929."
Line55: "Do."
Line62: "1st January, 1930. 1st January, 1933. 1st May, 1926."
Line73: "1st January, 1929."
Line79: "1st January, 1931."
Line85: "1st December, 1931."
That's 13 lines but some contain multiple dates. Total dates: line9:2, line10:1, line22:1, line28:2? "1927. Do." could be two: "1927" and "Do.", line33:1, line34:1, line37:2, line49:1, line55:1, line62:3, line73:1, line79:1, line85:1. Sum = 2+1+1+2+1+1+2+1+1+3+1+1+1 = 17? But we need 16. One extra maybe for something else.
But note: The Date of Appointment column might have one date per row. The "Date of First Appointment" is separate (bottom list). So we need 16 Date of Appointment values.
Let's assign sequentially to rows 1-16:
Row1: first date in line9: 25th February, 1929
Row2: second date in line9: 1st January, 1927
Row3: line10: Do. (so 1st January, 1927)
Row4: line18? But line18 is "Do." after row4 name. However, the dates might be in a separate column that OCR read after the names for each row? But the linear text order is not strictly row-by-row for all columns.
Given the salaries are in row order, perhaps the Date of Appointment, Authority, Allowances are also in row order but interleaved with other columns. Let's look at the lines between salaries.
Between row1 salary (line12) and row2 salary (line14): line13 is Authority for row2? Actually line11 is Authority for row1, line12 salary+allowance row1, line13 Authority row2, line14 salary row2. So Authority appears before salary for each row? For row1, Authority at line11, salary at line12. For row2, Authority at line13, salary at line14. For row3, Authority? Not seen before salary at line16. Line15 is "Do." which could be Authority for row3 (ditto row2). Then line16 salary row3. For row4, Authority? Not seen before salary at line20. Line17 name, line18 "Do.", line19 "Do.", line20 salary. The "Do." lines might be Office and Date of Appointment? But Authority missing? Could be "Do." from line15? But that's far.
For row5: line21 name, line22 date, line23 Authority, line24 salary. So row5 has Date, Authority, Salary.
Row6: line25 "Schools for New Territories," (Office), line26 "Do.", line27 name, line28 "1927. Do." (Date?), line29 "Do.", line30 salary. Authority not shown, maybe "Do." from row5?
Row7: line31 "Do.", line32 name, line33 "Do.", line34 "Do..", line35 name row8, line36 name row9, line37 date for row9? This is messy.
Given the complexity, but we have the bottom list for Date of First Appointment, and we have salaries, and we have some authorities and allowances. For the proofread text, we can reconstruct the table with the columns and fill in the data we are confident about, and use "..." for missing.
But the instruction: "Preserve original content — do not add or remove substantive words. Do not rephrase, rewrite, or change the author's style, tone, or word choice. Only correct what is clearly wrong." So we should output the text as a corrected version of the OCR, not a reconstructed database. However, the OCR is a table, so the corrected version should be a table.
The user said: "Return the proofread text in standard Markdown: Use Markdown table syntax to reconstruct tabular data."
Thus we need to produce a Markdown table that represents the original table.
Given the OCR is from a known source (Hong Kong Blue Book 1933?), we could try to replicate the exact table. But we must base on the OCR provided.
I think the best is to create a table with the columns as per header, and for each row, fill in the information we can extract from the OCR text, keeping the original wording (e.g., "Do." for ditto). We'll use the order of names as they appear.
Let's define columns:
We have 16 rows.
We'll populate row by row using the OCR data.
Row1:
Row2:
Row3:
Row4:
Row5:
Row6:
Row7:
Row8:
Row9:
Thus row9 starts at line42. So row9:
Row10:
Row11:
But in the linear text, after row10 salary (line51 4300), we have line52 "Class IV Clerk," line53 "( do.)" line54 "Leung Tak-wa," line55 "Do." line56 "( do. )" line57 "Ho Ka-chun." line58 "(1)" line59 "Class V Clerk," line60 "( do.)" line61 "Chan Hon-shan." line62 "1st January, 1930. 1st January, 1933. 1st May, 1926." line63 "C.S.O. 5254 of 1930," line64 "1,800 $60 Shorthand Allowance," line65 "C.S.O. 5254 of 1933," line66 "1,500" line67 "C.S.O. 3646 of 1926." line68 "1,400" line69 "I" line70 "Do," line71 "( do.)" line72 "Luk Hok-king." line73 "1st January, 1929." line74 "******" line75 "1,400" line76 "Do.," line77 "( do. )" line78 "To Man-to." line79 "1st January, 1931." line80 "C.S.O. 5255 of 1931." line81 "1,300" line82 "Do.," line83 "( do.)" line84 "Chan Sik-pui," line85 "1st December, 1931." line86 "C.S.O. 5255 of 1931." line87 "1,300"
So the salaries for row11 (Leung Tak-wa) and row12 (Ho Ka-chun) are not explicitly shown as numbers in this segment? But the salary list says they are 1800 and 1500. However, line64 is 1800 for Chan Hon-shan (row13). Line66 is 1500 for row14? But row14 is Luk Hok-king? Actually row13 Chan Hon-shan, row14 Luk Hok-king, row15 To Man-to, row16 Chan Sik-pui. The salaries: row13 1800, row14 1500, row15 1400, row16 1400? But the list says row13 1400, row14 1400, row15 1300, row16 1300. There's a mismatch.
Let's map the linear text salaries to rows based on the order of names in the linear text. The names appear in this order in linear text:
The salaries appear at lines: 12,14,16,20,24,30,39,41,45,51,64,66,68,75,81,87. That's 16 salaries. Let's match by index:
Index 1: line12 -> 5400 (row1)
2: line14 -> 3800 (row2)
3: line16 -> 3200 (row3)
4: line20 -> 3200 (row4)
5: line24 -> 2550 (row5)
6: line30 -> 2400 (row6)
7: line39 -> 1950 (row7)
8: line41 -> 1950 (row8)
9: line45 -> 1650 (row9)
10: line51 -> 4300 (row10)
11: line64 -> 1800 (row11)
12: line66 -> 1500 (row12)
13: line68 -> 1400 (row13)
14: line75 -> 1400 (row14)
15: line81 -> 1300 (row15)
16: line87 -> 1300 (row16)
But row11 is Leung Tak-wa, row12 Ho Ka-chun, row13 Chan Hon-shan, row14 Luk Hok-king, row15 To Man-to, row16 Chan Sik-pui. So salaries: Leung Tak-wa 1800, Ho Ka-chun 1500, Chan Hon-shan 1400, Luk Hok-king 1400, To Man-to 1300, Chan Sik-pui 1300. That seems plausible: Class IV Clerks might have higher salary than Class V? But here Leung Tak-wa (Class IV) gets 1800, Ho Ka-chun (Class IV) gets 1500? That's odd. But maybe Ho Ka-chun is Class V? The text says "Class IV Clerk, ( do.) Leung Tak-wa, Do. ( do. ) Ho Ka-chun. (1) Class V Clerk, ( do.) Chan Hon-shan." So Ho Ka-chun is also Class IV? But then salary 1500 for Class IV? Could be. Then Chan Hon-shan Class V 1400, Luk Hok-king Class V 1400, To Man-to Class V 1300, Chan Sik-pui Class V 1300. That matches the salary list.
Thus the linear text salaries are in correct row order.
Now, for each row, we need to extract Date of Appointment, Authority, Allowances from the linear
Date
OFFICE,
NAME.
of Appointment.
Authority.
Auunal Salary.
EDUCATION DEPARTMENT,—Continued.
House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose.
Leave taken during 1933.
Inspector of Vernacular Schools, Liang Sai-wa, B.A.
Sub-Inspector of Vernacular
Lau Shuk-chong.
Schools, Urban District,
Do.,
Do..
Sub-Inspector of Vernacular
Lam Tun-wing.
25th February, 1929. 1st January, 1927.
Do.
C.S.O. 5006 of 1929.
$5,400 $25 from the Technical Institute.
C.S.O. 888 of 1926.
3,800
Do.
3,200
Chan Tni-shang,
Do.
Do.
3,200
Lai Kuk-yoen.
17th February,
C.S.O. 6502 of 1910,
2,550
Schools for New Territories,
Do.,
Hui Tat-cheung.
Do.
2,400
Do.,
Chan Kam-tong.
Do.,
Do..
Chan Yuk-cheung.
To Chung-pok.
1st January, 1927. Do.
Do.
1,950
Do.
1,950
2nd April,
1928.
C.S.O. 4841 of 1928.
1,650
Class I Clerk,
(J. C. S.)
S. M. Suddick.
1st January, 1929.
C.S.O. 5254 of 1929.
4,300
Class IV Clerk,
( do.)
Leung Tak-wa,
Do.
( do. )
Ho Ka-chun.
(1)
Class V Clerk,
( do.)
Chan Hon-shan.
1st January, 1930. 1st January, 1933. 1st May, 1926.
C.S.O. 5254 of 1930,
1,800 $60 Shorthand Allowance,
C.S.O. 5254 of 1933,
1,500
C.S.O. 3646 of 1926.
1,400
I
Do,
( do.)
Luk Hok-king.
1st January, 1929.
******
1,400
Do.,
( do. )
To Man-to.
1st January, 1931.
C.S.O. 5255 of 1931.
1,300
Do.,
( do.)
Chan Sik-pui,
1st December, 1931.
C.S.O. 5255 of 1931.
1,300
****
Date of First Appointment.
ART
274
(J 144 )
25th February,
1st March. 1923.
Do.
2nd April, 1928. 10th April, 1903. 20th June, 1922.
7th January,
1923.
1st September, 1926.
No comments yet.
Private notes are available after approval.