The user wants me to proofread OCR output of a historical record relating to Hong Kong, specifically a table of staff for the Kowloon-Canton Railway Traffic Staff. The OCR text is messy with broken lines, misaligned columns, and missing structure. I need to reconstruct it as a proper Markdown table.
First, I need to understand the columns. The header row shows:
OFFICE. NAME. Date of Authority. Annual Salary. Appointment. House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose. Leave taken during 1933. Date of First Appointment.
But the OCR has "302 (J 172 )" which might be page number and reference. Then the table starts with "Ticket Collector," and then names.
Looking at the data rows, each entry seems to have:
But the OCR is jumbled. Let me parse line by line.
The text after header:
"302 (J 172 ) Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929. Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929. Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928. Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930. Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930. Do Do. 510 Do. 5 days. 2 Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932. Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933. Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933. Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932. Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days. Do.. Chui Che-kwong. 1st August, Do. 420 1933. Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933. Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923. Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920. Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
This is a mess. It appears the OCR has merged multiple rows and columns. The "Do." likely means "Ditto" (same as above). The numbers (1), (2), etc. might be row numbers.
I need to reconstruct the table. Let's identify each staff entry.
The header: "KOWLOON-CANTON RAILWAY,—Continued, TRAFFIC STAFF, — Continued." So this is a continuation of a table.
Columns (from header):
But the OCR shows "Date of Authority." and "Annual Salary." and "Appointment." maybe "Date of Appointment"? Actually header: "Date of Authority. Annual Salary. Appointment." Wait the header line: "OFFICE. NAME. Date of Authority. Annual Salary. Appointment. House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose. Leave taken during 1933. Date of First Appointment."
But there is also "Date of First Appointment" at the end. And "Appointment" might be "Date of Appointment"? Actually the header says "Date of Authority. Annual Salary. Appointment." Could be "Date of Appointment"? But the OCR shows "Date of Authority." Might be "Date of Appointment"? Let's examine data: For Luk Kim-bung: "9th January, Establishment. $510 1929." That seems like Date of Appointment: 9th January 1929? But "Establishment" is Authority. Then "1929" might be year of appointment? But then "Date of First Appointment" column shows "1st October 1922" for first entry? Actually at the end there is a list of dates: "1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933." Those might be the "Date of First Appointment" for each staff.
Let's map each staff.
Staff list:
But the OCR shows "Brakesman, Leung Hin." then "Tong Cheong." then "Teang Yan." Might be three brakesmen.
Now, each row should have 8 columns. Let's try to parse each.
First row: "Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929." Then "Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929." This suggests that "Do." means ditto for Office (Ticket Collector) and Authority (Establishment). So each row: Office, Name, Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken, Date of First Appointment.
But the "Date of First Appointment" appears at the end in a block. Possibly the OCR separated that column into a list at the bottom. That is common in OCR of tables: the last column gets read as a block.
Thus we need to match each staff with a date from that list. There are 17 staff? Let's count the dates in the block:
But the block at the end includes: "Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923." That seems like leave dates? Actually "Leave taken during 1933" column might have multiple dates. The block might be a mix of leave dates and first appointment dates.
This is tricky. Perhaps the original table had two sections: Ticket Collectors and Brakesmen. The "Date of First Appointment" column might be the last column. The OCR has dumped all the dates from that column at the end.
Let's look at the structure: The header includes "Date of First Appointment." at the end. In the data, after each row, there might be a date. But the OCR shows "1929." at end of first row, then "1929." at end of second row, etc. Actually first row ends with "1929." second row ends with "1929." third row ends with "1928." fourth row ends with "1930." fifth row ends with "1930." then "1930." again? Then "1932." then "1922." then "1933." then "1932." then "1933." then "1932." then "1933." then "1933." then "1923." then "1920." then "1922." then "1932." then "1933." then "1933." then "1933." then "1912." then "1918." then "1922." then "1923."
But the block at the end seems to be a list of dates that might correspond to "Date of First Appointment" for each staff in order.
Let's list staff in order as they appear:
That's 17 staff.
Now the block of dates at the end (after "Teang Yan. 1st March, Do. 600 10 days.") there is "Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923." That's 10 items? Actually count:
That's 9 items. But there are 17 staff. So not matching.
Maybe the "Date of First Appointment" column is actually the "Date of Appointment" (the third column) and the last column is "Date of First Appointment" which might be the same as the third column for some? But the header has both "Date of Authority" and "Date of First Appointment". Actually header: "Date of Authority. Annual Salary. Appointment. House or Quarters... Leave taken during 1933. Date of First Appointment." Wait there are 8 columns: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1933, Date of First Appointment.
But "Appointment" might be "Date of Appointment"? Or "Authority"? The header says "Date of Authority." then "Annual Salary." then "Appointment." That seems like three columns: Date of Authority, Annual Salary, Appointment. But "Appointment" could be "Date of Appointment"? However the data shows "Establishment" under Authority? For first row: "9th January, Establishment. $510 1929." So "9th January" might be Date of Authority? "Establishment" is Authority. "$510" is Annual Salary. "1929" might be Appointment year? But then "Date of First Appointment" is separate.
Let's examine second row: "1st December, Do. 540 Quarters. Do. ..... 10 days. 1929." So "1st December" Date of Authority, "Do." (Establishment) Authority, "540" Salary, "Quarters" House/Quarters, "Do." (maybe ditto for Allowances?), "....." maybe leave? "10 days" leave, "1929" Date of First Appointment? But then the block at the end has "1st October 1922" for first staff? That doesn't match.
Maybe the "1929" at end of first row is actually the Date of First Appointment for Luk Kim-bung? But the block at the end starts with "1st October 1922". That could be for Luk Kim-bung? But he started 1929? Hmm.
Let's search for known historical records. This looks like the "Hong Kong Blue Book" or "Civil List" for 1933. The Kowloon-Canton Railway staff list. I recall that the Blue Book has tables with columns: Office, Name, Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during year, Date of First Appointment.
In many Blue Books, the columns are: Office, Name, Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933, Date of First Appointment.
The header in OCR: "OFFICE. NAME. Date of Authority. Annual Salary. Appointment. House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose. Leave taken during 1933. Date of First Appointment."
But "Date of Authority" might be a mis-OCR for "Date of Appointment". And "Appointment" might be "Date of Appointment"? Actually there are two date columns: Date of Appointment and Date of First Appointment. The header shows "Date of Authority." and "Appointment." Could be "Date of Appointment" and "Date of First Appointment"? But then "Authority" is separate? Wait typical Blue Book columns: "Office", "Name", "Date of Appointment", "Authority", "Annual Salary", "House/Quarters/Allowances", "Leave taken during 1933", "Date of First Appointment". That's 8 columns. The OCR header has 8 items: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House or Quarters..., Leave taken during 1933, Date of First Appointment. So "Date of Authority" might be "Date of Appointment", "Appointment" might be "Authority"? But "Authority" is usually the appointing authority (e.g., "Establishment"). In the data, "Establishment" appears after the date. So column 3: Date of Appointment, column 4: Authority, column 5: Annual Salary, column 6: House/Quarters/Allowances, column 7: Leave taken, column 8: Date of First Appointment.
But the header says "Date of Authority. Annual Salary. Appointment." That's three items. Could be mis-ordered. Let's assume the correct columns are:
Now parse each row accordingly.
Row 1: "Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929."
Row 2: "Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929."
Row 3: "Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928."
Row 4: "Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930."
Row 5: "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930."
This is messy. "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930."
Probably two rows: Tang K-wing and Li Hon-ki.
Let's continue reading sequentially after "1930." The text: "Do Do. 510 Do. 5 days. 2 Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932. Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933. Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933. Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932. Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days. Do.. Chui Che-kwong. 1st August, Do. 420 1933. Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933. Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923. Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920. Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
This is a run-on. It seems the OCR has lost line breaks. We need to split into rows.
Let's identify each staff entry by the pattern: "Do., [Name]." or "Do.. [Name]." or "Brakesman, [Name]."
List of names in order after Tang K-wing:
That's 12 more, total 17 (including first 5). Good.
Now for each, we need to extract the fields.
The pattern for each row (after the first few) seems to be:
"Do., [Name]. [Date of Appointment], Do. [Salary] [House/Quarters/Allowances] [Leave] [Date of First Appointment?]"
But the "Date of First Appointment" might be the last date in each row. However, the block of dates at the end might be the "Date of First Appointment" for all staff, but they are jumbled.
Let's parse each row from the text sequentially, using the fact that "Do." means ditto for Office and Authority.
We'll go through the text after "Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930."
Actually, let's copy the text from "Do Do. 510 Do. 5 days. 2" onward.
Full text after "1930." (which might be end of Tang K-wing row):
"Do Do. 510 Do. 5 days. 2 Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932. Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933. Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933. Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932. Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days. Do.. Chui Che-kwong. 1st August, Do. 420 1933. Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933. Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923. Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920. Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
We need to split into rows. Each row starts with "Do.," or "Do.." or "Brakesman,".
Let's split by "Do.," and "Do.." and "Brakesman," but careful.
Better to manually parse.
We'll create a list of rows with fields.
First, the header row: we'll create a Markdown table with columns: Office, Name, Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933, Date of First Appointment.
Now, row 1: Luk Kim-bung
Let's look at the block of dates after Pun Yuk-shui: "1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933." That's 9 dates. There are 9 Ticket Collectors before Li Hung-chuen? Actually we have 14 Ticket Collectors total. But the block has 9 dates. Then later "5th January, 1932. 1933. 4th February, 1932. 1st August, 1933. 13th October, 1933. 1923. 1920. 1922." That's more.
Maybe the "Date of First Appointment" is actually the last column in each row, and the OCR has placed them inline. For example, for Luk Kim-bung, the row ends with "1929." which could be his Date of First Appointment. For Ng Kau, ends with "1929." For Au Ngok-sang, ends with "1928." For Lam Ping-ki, ends with "1930." For Tang K-wing, ends with "1930." For Li Hon-ki, the text "1930." appears after his name? Actually "Li Hon-ki, 1930." Then "Do Do. 510 Do. 5 days. 2" might be his row? Let's parse Li Hon-ki row.
After "Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930." Then "Do Do. 510 Do. 5 days. 2" This might be the continuation of Li Hon-ki's row: "Do Do." means Office and Authority ditto. "510" salary. "Do." House/Quarters ditto (Quarters). "5 days" leave. "2" ??? Maybe "2" is a typo or part of next.
Then "Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932." So Cheung Chung-hong row: Date of Appointment: 17th February, Authority: Establishment, Salary: 450, House/Quarters: Do. (Quarters), Leave: 7 days, Date of First Appointment: 1932.
Then "Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922." Wait "1st October 1922" appears here. That might be Pun Yuk-shui's Date of First Appointment. But then the block continues with more dates. Actually after "10 days." we have "1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933." That looks like a list of Date of First Appointment for the previous 9 Ticket Collectors? But they are placed after Pun Yuk-shui. Possibly the OCR mis-ordered: the Date of First Appointment column for all Ticket Collectors got dumped after Pun Yuk-shui.
Let's check the original Blue Book layout. In such tables, the Date of First Appointment column is at the right. The OCR might have read the table column by column? But the text seems to read row by row, but the last column values are all placed at the end of the Ticket Collector section.
Given the difficulty, perhaps the best is to reconstruct the table as the OCR presents it, but with proper row separation and column alignment based on the apparent pattern. Since the user wants proofreading of OCR output, we should correct the text flow and present it as a clean table. We need to infer the correct columns for each row.
We have 17 staff. We'll create a table with 8 columns. We'll fill in the data as per the OCR text, using the dates that appear in each row for the first 7 columns, and for the 8th column (Date of First Appointment), we might need to use the block dates in order.
But the block dates: first block after Pun Yuk-shui has 9 dates. There are 9 Ticket Collectors before Li Hung-chuen? Let's count Ticket Collectors up to Pun Yuk-shui:
That's 8. But the block has 9 dates. Maybe includes Li Hung-chuen? But Li Hung-chuen appears later.
Second block: after "1933." we have "Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933." Then "Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932." Then "Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days." Then "Do.. Chui Che-kwong. 1st August, Do. 420 1933." Then "Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933." Then Brakesmen.
The dates at the very end: "Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923." These might be leave dates for the Brakesmen? Or Date of First Appointment for Brakesmen.
Given the complexity, perhaps the original table had the Date of First Appointment column filled for each row, but the OCR run-on merged them. The proofreading task is to fix the text flow and present the table correctly. Since we cannot know the exact original, we must do our best to reconstruct a plausible table.
Given the instructions: "Reconstruct Sentences and Paragraphs" and "Tables — if data is clearly in table format, reconstruct the table using Markdown table syntax." We should output a Markdown table with the data as it appears, but with rows properly separated and columns aligned.
We need to decide on the columns. The header in the OCR: "OFFICE. NAME. Date of Authority. Annual Salary. Appointment. House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose. Leave taken during 1933. Date of First Appointment."
But "Date of Authority" and "Appointment" are separate. In the data, we see "9th January, Establishment." So "9th January" is likely "Date of Appointment" and "Establishment" is "Authority". Then "Annual Salary" is $510. Then "Appointment" column? That might be "Date of Appointment" again? No.
Maybe the columns are: Office, Name, Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933, Date of First Appointment. That's 8 columns. The header has 8 items if we consider "Date of Authority" as "Date of Appointment", "Annual Salary", "Appointment" as "Authority"? But "Appointment" is singular. Actually the header: "Date of Authority. Annual Salary. Appointment." That's three. But there are 8 columns total. Let's count header items: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House or Quarters..., Leave taken during 1933, Date of First Appointment. That's 8. So "Date of Authority" is column 3, "Annual Salary" column 4, "Appointment" column 5, "House or Quarters..." column 6, "Leave taken during 1933" column 7, "Date of First Appointment" column 8.
But in the data, after Name, we have a date, then "Establishment", then salary, then house/quarters, then leave, then date. That would make "Establishment" the "Appointment" column? But "Appointment" might be the appointing authority. In Hong Kong Blue Books, the column is "Authority" (e.g., Establishment, Secretary of State, etc.). So "Appointment" might be a mis-OCR for "Authority". And "Date of Authority" might be "Date of Appointment". So columns: Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken, Date of First Appointment. That's 6 columns after Name. Plus Office and Name = 8. That matches.
Thus the header should be interpreted as:
We'll adopt that.
Now, for each row, we need to extract these 8 fields.
Let's parse row by row from the OCR text, assuming the text flows row by row but with some missing line breaks.
We'll go through the text sequentially and split into rows when we see a new name with "Do.," or "Brakesman,".
First, the text starts with "302 (J 172 )" which is likely page number and reference. We'll keep that as a header note.
Then "Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929."
This is row 1. Fields:
Let's check row 2: "Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929."
Here, "1929." at end. So Date of First Appointment = 1929.
Row 3: "Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928."
Ends with 1928.
Row 4: "Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930."
Ends with 1930.
Row 5: "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930."
This seems to have two rows merged. "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.." Then "Li Hon-ki, 1930." So Tang K-wing row ends with "Do." (maybe leave?) and then "Li Hon-ki, 1930." starts next row. But Tang K-wing's Date of First Appointment might be 1930? The "1930." after Li Hon-ki might be for Li Hon-ki.
Then "Do Do. 510 Do. 5 days. 2" This might be Li Hon-ki's row: "Do Do." (Office and Authority), "510" salary, "Do." house, "5 days" leave, "2" ??? Maybe "2" is a typo for "1932"? But then next row "Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932." So Cheung Chung-hong ends with 1932.
Then "Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922." Here we have a full date "1st October 1922" at end. That could be Pun Yuk-shui's Date of First Appointment.
Then a list of dates: "5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933." These might be the Date of First Appointment for the previous 8 Ticket Collectors (Luk Kim-bung to Pun Yuk-shui) but in reverse order? Let's see: If we assign them in order of rows:
But Pun Yuk-shui already has "1st October 1922" inline. That conflicts.
Maybe the inline "1st October 1922" is actually the Date of First Appointment for Luk Kim-bung, and the list continues for the rest. But the list starts with "5th January, 1928" after "1st October 1922." So perhaps the Date of First Appointment for each Ticket Collector is given in that block, and the inline "1929." etc. are actually the year of Date of Appointment? Let's examine.
In row 1, "9th January, Establishment. $510 1929." The "1929" could be the year of Date of Appointment (9th January 1929). Then the Date of First Appointment is not given inline, but appears in the block. In row 2, "1st December, Do. 540 Quarters. Do. ..... 10 days. 1929." The "1929" could be Date of Appointment year (1st December 1929). Row 3: "1st November, Do. 540 $60 Rent Allowance. 1928." -> 1st November 1928. Row 4: "11th July, Do. 480 Quarters. 1930." -> 11th July 1930. Row 5: "1st September, Do. 480 Do. Do.." -> 1st September 1930? Then Li Hon-ki: "Li Hon-ki, 1930." -> maybe 1930 is his Date of Appointment year. Then "Do Do. 510 Do. 5 days. 2" -> the "2" might be a typo for "1932"? But Cheung Chung-hong has "1932" at end. Pun Yuk-shui has "1st October 1922" at end? That doesn't fit.
Alternatively, the "1929." at end of row 1 might be the Date of First Appointment. But then the block of dates would be extraneous.
Given the confusion, I think the best approach for proofreading is to present the table as the OCR text suggests, with each row having the fields as they appear in sequence, and for the Date of First Appointment, use the value that appears at the end of each row in the OCR (like 1929, 1929, 1928, 1930, 1930, 1930, 1932, 1st October 1922, etc.) and for rows where it's missing, leave blank or use the block dates in order.
But the user expects a clean Markdown table. We need to make a decision.
Let's look at the very end of the OCR: "Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923." These might be the Leave taken during 1933 for the Brakesmen? Or Date of First Appointment for Brakesmen.
For Brakesmen:
This is too messy. Perhaps the OCR has completely garbled the last part.
Given the instructions, we are to proofread the OCR output, not to perfectly reconstruct the original table if it's ambiguous. We should correct obvious OCR errors, rejoin broken lines, and present the text in a readable format. For tables, we should reconstruct the table as best as we can.
I think the best is to output the table with the rows as they appear in the OCR, using the columns as per the header, and for each row, fill the columns with the data that appears in the text in order. We'll need to split the run-on text into rows.
Let's manually split the text into rows based on the pattern "Do.," "Do.." "Brakesman,".
We'll write a script in mind, but let's do it manually.
Full text after header:
"302 (J 172 ) Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929. Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929. Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928. Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930. Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930. Do Do. 510 Do. 5 days. 2 Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932. Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933. Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933. Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932. Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days. Do.. Chui Che-kwong. 1st August, Do. 420 1933. Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933. Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923. Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920. Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
We need to identify each row. The rows are separated by "Do.," or "Do.." or "Brakesman,". But note that "Do.." appears sometimes as "Do.." and sometimes as "Do.".
Let's split by the pattern: (^|\\s)(Do\\.\\.?|Brakesman), but careful.
Better to list the names in order as they appear with their preceding "Do." or "Brakesman".
That's 18 entries? But we have 17 staff. Li Ping-sum might be separate. Let's count: 1-5 = 5, 6 Li Hon-ki, 7 Cheung Chung-hong, 8 Pun Yuk-shui, 9 Li Hung-chuen, 10 Cheng Wei-cliu, 11 Ho Kwong-wai, 12 Chui Che-kwong, 13 Leung Kam-to, 14 Li Ping-sum, 15 Leung Hin, 16 Tong Cheong, 17 Teang Yan. That's 17. Good.
Now for each, we need to extract the fields. The fields after the name are: Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken, Date of First Appointment.
But the text for each row is not clearly delimited. We'll have to parse the text between each name.
Let's write the text as a sequence of tokens.
I'll copy the text into a string and simulate parsing.
Text: "302 (J 172 ) Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929. Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929. Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928. Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930. Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930. Do Do. 510 Do. 5 days. 2 Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932. Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933. Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933. Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932. Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days. Do.. Chui Che-kwong. 1st August, Do. 420 1933. Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933. Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923. Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920. Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
We'll parse sequentially.
We'll maintain current office = "Ticket Collector" for first many rows, then "Brakesman" for last three.
Authority seems to be "Establishment" for all (Do. means ditto).
Let's parse row by row.
Row 1: "Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929."
Row 2: "Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929."
Row 3: "Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928."
Row 4: "Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930."
Row 5: "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930."
This is two rows merged. Let's split at "Do.. Li Hon-ki". So Row 5a: Tang K-wing.
Row 5b: Li Hon-ki: "Do.. Li Hon-ki, 1930. Do Do. 510 Do. 5 days. 2"
Row 6: Cheung Chung-hong: "Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932."
Row 7: Pun Yuk-shui: "Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922."
Then we have a list of dates: "5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933." These are 8 dates. They might be the Date of First Appointment for the previous 8 Ticket Collectors (Luk Kim-bung to Pun Yuk-shui) but in reverse order? Let's see: If we assign in order of rows 1-7 plus maybe Li Hon-ki? That's 8 rows (1 Luk, 2 Ng, 3 Au, 4 Lam, 5 Tang, 6 Li Hon-ki, 7 Cheung, 8 Pun). The dates list has 8 dates. The first date in list is "5th January, 1928" but Pun Yuk-shui already has "1st October 1922". So maybe the list corresponds to rows 1-7 and Li Hon-ki? But Pun Yuk-shui is row 8. The list appears after Pun Yuk-shui. Could be that the Date of First Appointment for rows 1-7 are given in this list, and Pun Yuk-shui's is given inline. But then the list has 8 dates, rows 1-7 are 7 rows. Li Hon-ki is row 6? Actually row 6 is Cheung, row 7 is Pun. Li Hon-ki is row 5b. So rows 1-5a, 5b, 6, 7 = 8 rows. The list has 8 dates. Let's map:
Row 1 Luk Kim-bung: 5th January, 1928? But the list starts with "5th January, 1928." Then "20th December, 1927." etc. But the first date in the list is "5th January, 1928." However, the block starts with "1st October 1922." which is attached to Pun Yuk-shui. Then "5th January, 1928." So the list after Pun Yuk-shui includes "5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933." That's 8 dates. If we consider that the Date of First Appointment for the 8 Ticket Collectors (Luk, Ng, Au, Lam, Tang, Li Hon-ki, Cheung, Pun) are given in that order? But Pun already has a date. Maybe the list is for the first 7 and Li Hung-chuen? But Li Hung-chuen comes later.
This is too speculative. For proofreading, we might keep the inline dates as they appear in each row, and for the block, we might note them as a separate list? But the instruction says to reconstruct the table. Perhaps the original table had the Date of First Appointment column, and the OCR read it as a block at the end of the Ticket Collector section. In that case, we should distribute those dates to the rows in order.
Given the difficulty, I'll assume that the Date of First Appointment for each Ticket Collector is given in the block in the same order as the rows. The block starts with "1st October 1922." (for Luk Kim-bung), then "5th January, 1928." (for Ng Kau), "20th December, 1927." (for Au Ngok-sang), "1st November, 1925." (for Lam Ping-ki), "3rd August, 1930." (for Tang K-wing), "11th July, 1930." (for Li Hon-ki), "1st September, 1930." (for Cheung Chung-hong), "1st May, 1930." (for Pun Yuk-shui), "1933." (for Li Hung-chuen). But there are 9 dates for 9 Ticket Collectors up to Li Hung-chuen? Actually we have 14 Ticket Collectors. The block has 9 dates. Then later more dates appear for the remaining.
Let's see later: After "1933." we have "Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933." So Li Hung-chuen has "5th January, 1932" and "1933" maybe Date of First Appointment? Then "Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932." So Cheng Wei-cliu has "4th February, 1932" as Date of First Appointment? Then "Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days." No date. Then "Do.. Chui Che-kwong. 1st August, Do. 420 1933." So Chui Che-kwong has "1933". Then "Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933." This is messy.
Given the time, I think the best is to output a cleaned-up version of the text as a table with the columns as per header, and for each row, include the data that appears in the OCR in the order they appear, leaving blank where missing. We'll not try to guess the missing Date of First Appointment from the block; instead, we'll include the block as a separate note? But the instruction says to reconstruct the table. Maybe we can create two tables: one for Ticket Collectors, one for Brakesmen, with the columns.
Given the user is a professional editor, they might expect a clean table with all rows and columns filled as per the original. But we don't have the original. We only have OCR. The task is to proofread the OCR output, i.e., correct the OCR errors and format it nicely.
Thus, we should produce a Markdown document that represents the corrected text. The text includes a header "KOWLOON-CANTON RAILWAY,—Continued, TRAFFIC STAFF, — Continued." Then a table. We'll output the table in Markdown.
We'll define the columns as: Office, Name, Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933, Date of First Appointment.
We'll parse each row as best we can.
Let's do it systematically by writing a small parser in mind.
We'll split the text into rows by the delimiter "Do.," "Do.." "Brakesman," but also the first row starts with "Ticket Collector,".
We'll use the following approach: The text after "302 (J 172 )" is a sequence of records. Each record starts with either "Ticket Collector," or "Do.," or "Do.." or "Brakesman,". The record ends before the next such start.
We'll split the string by the regex /(?:^|\s)(Do\.\.?|Brakesman),/ but keep the delimiter.
Let's manually split.
I'll write the text with markers.
Start: "302 (J 172 ) Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929. Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929. Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928. Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930. Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.. Li Hon-ki, 1930. Do Do. 510 Do. 5 days. 2 Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932. Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933. Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933. Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932. Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days. Do.. Chui Che-kwong. 1st August, Do. 420 1933. Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do... Li Ping-sum. 1933. Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923. Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920. Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
Now, find each occurrence of "Do..", "Do.,", "Do Li", "Do...", "Brakesman,".
Positions:
So 18 starts, but some are continuations. The "Do Do." and "Do... Li Ping-sum" might be part of previous rows.
Let's assign each staff to a start:
That's 17 staff. Good.
Now for each, we need to extract the text until the next start.
We'll define the text for each row as from the start keyword to just before the next start keyword.
Let's do that.
Row 1 (Luk Kim-bung): from "Ticket Collector, Luk Kim-bung." to before "Do.. Ng Kau."
Text: "Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929."
Row 2 (Ng Kau): from "Do.. Ng Kau." to before "Do., Au Ngok-sang."
Text: "Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929."
Row 3 (Au Ngok-sang): from "Do., Au Ngok-sang." to before "Do., Lam Ping-ki."
Text: "Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928."
Row 4 (Lam Ping-ki): from "Do., Lam Ping-ki." to before "Do., Tang K-wing."
Text: "Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930."
Row 5 (Tang K-wing): from "Do., Tang K-wing." to before "Do.. Li Hon-ki,"
Text: "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.." Wait the next start is "Do.. Li Hon-ki," but the text includes "Do.." at end? Actually the text before "Do.. Li Hon-ki," is "Do., Tang K-wing. (5) 1st September, Do. 480 Do. Do.." So the "Do.." at end might be the start of next row. So Row 5 text: "Do., Tang K-wing. (5) 1st September, Do. 480 Do." (excluding the trailing "Do..")
Row 6 (Li Hon-ki): from "Do.. Li Hon-ki," to before "Do., Cheung Chung-hong."
Text: "Do.. Li Hon-ki, 1930. Do Do. 510 Do. 5 days. 2"
Row 7 (Cheung Chung-hong): from "Do., Cheung Chung-hong." to before "Do., Pun Yuk-shui."
Text: "Do., Cheung Chung-hong. 17th February, Do. 450 Do. 7 days. 1932."
Row 8 (Pun Yuk-shui): from "Do., Pun Yuk-shui." to before "Do Li Hung-chuen,"
Text: "Do., Pun Yuk-shui. 3rd January, Do. 420 Do. 10 days. 1st October 1922. 5th January, 1928. 20th December, 1927. 1st November, 1925. 3rd August, 1930. 11th July, 1930. 1st September, 1930. 1st May, 1930. 1933."
Row 9 (Li Hung-chuen): from "Do Li Hung-chuen," to before "Do., Cheng Wei-cliu."
Text: "Do Li Hung-chuen, 7th April, Do. 420 5 days. 5th January, 1932. 1933."
Row 10 (Cheng Wei-cliu): from "Do., Cheng Wei-cliu." to before "Do.. Ho Kwong-wai,"
Text: "Do., Cheng Wei-cliu. Do. Do. 420 9 days. 4th February, 1932."
Row 11 (Ho Kwong-wai): from "Do.. Ho Kwong-wai," to before "Do.. Chui Che-kwong."
Text: "Do.. Ho Kwong-wai, Do. Do. 420 Quarters. 10 days."
Row 12 (Chui Che-kwong): from "Do.. Chui Che-kwong." to before "Do., Leung Kam-to."
Text: "Do.. Chui Che-kwong. 1st August, Do. 420 1933."
Row 13 (Leung Kam-to): from "Do., Leung Kam-to." to before "Do... Li Ping-sum."
Text: "Do., Leung Kam-to. Do. Do. 420 Quarters. 13th October, Do. 420 Do..."
Row 14 (Li Ping-sum): from "Do... Li Ping-sum." to before "Brakesman, Leung Hin."
Text: "Do... Li Ping-sum. 1933."
Row 15 (Leung Hin): from "Brakesman, Leung Hin." to before "Do., Tong Cheong."
Text: "Brakesman, Leung Hin. 1st July, Do. 600 10 days. 1923."
Row 16 (Tong Cheong): from "Do., Tong Cheong." to before "Do.. Teang Yan."
Text: "Do., Tong Cheong. 1st April, Do. 600 Quarters. 7 days. 1920."
Row 17 (Teang Yan): from "Do.. Teang Yan." to end.
Text: "Do.. Teang Yan. 1st March, Do. 600 10 days. Ist March, 1932. 9th January, 1933. 8th February, 1933. 4th April, 1933. 1st February, 1st March, 1912. 1st June, 1918. 1922. 1923."
Now, for each row, parse the fields.
We'll assume the fields after the name are in order: Date of Appointment, Authority, Annual Salary, House/Quarters/Allowances, Leave taken during 1933, Date of First Appointment.
But the number of fields varies. We'll parse by splitting on periods and commas? The text uses periods to separate fields? Not consistent.
Better to parse by known patterns: The Date of Appointment is a date like "9th January", "1st December", etc. Then "Establishment" or "Do." for Authority. Then a number for Salary. Then a string for House/Quarters/Allowances (could be "Quarters", "$60 Rent Allowance", "Do."). Then leave: "10 days", "5 days", "..... 10 days", "7 days", etc. Then a date for Date of First Appointment.
But the text often runs together.
Let's parse each row manually.
Row 1: "Ticket Collector, Luk Kim-bung. (1) 9th January, Establishment. $510 1929."
Row 2: "Do.. Ng Kau. (2) 1st December, Do. 540 Quarters. Do. ..... 10 days. 1929."
We'll set House/Quarters/Allowances = Quarters, Leave = 10 days, Date of First Appointment = 1929.
Row 3: "Do., Au Ngok-sang. (3) 1st November, Do. 540 $60 Rent Allowance. 1928."
Row 4: "Do., Lam Ping-ki. (4) 11th July, Do. 480 Quarters. 1930."
Row
OFFICE.
NAME.
Date of
Authority.
Annual Salary.
Appointment.
KOWLOON-CANTON RAILWAY,—Continued,
TRAFFIC STAFF, — Continued.
House or Quarters, and Allowances
for Rent, Entertainment, Personal, or for any other purpose.
Leave taken during 1933.
Date of First Appointment.
302
(J 172 )
Ticket Collector,
Luk Kim-bung.
(1)
9th January,
Establishment.
$510
1929.
Do..
Ng Kau.
(2)
1st December,
Do.
540
Quarters.
Do.
.....
10 days.
1929.
Do.,
Au Ngok-sang.
(3)
1st November,
Do.
540 $60 Rent Allowance.
1928.
Do.,
Lam Ping-ki.
(4)
11th July,
Do.
480
Quarters.
1930.
Do.,
Tang K-wing.
(5)
1st September,
Do.
480
Do.
Do..
Li Hon-ki,
1930.
Do
Do.
510
Do.
5 days.
2
Do.,
Cheung Chung-hong.
17th February,
Do.
450
Do.
7 days.
1932.
Do.,
Pun Yuk-shui.
3rd January,
Do.
420
Do.
10 days.
1st October
1922.
5th January,
1930.
1st May,
1930.
1933.
Do
Li Hung-chuen,
7th April,
Do.
420
5 days.
5th January,
1932.
1933.
Do.,
Cheng Wei-cliu.
Do.
Do.
420
9 days.
4th February,
1932.
Do..
Ho Kwong-wai,
Do.
Do.
420
Quarters.
10 days.
Do..
Chui Che-kwong.
1st August,
Do.
420
1933.
Do.,
Leung Kam-to.
Do.
Do.
420
Quarters.
13th October,
Do.
420
Do...
Li Ping-sum.
1933.
Brakesman,
Leung Hin.
1st July,
Do.
600
10 days.
1923.
Do.,
Tong Cheong.
1st April,
Do.
600
Quarters.
7 days.
1920.
Do..
Teang Yan.
1st March,
Do.
600
10 days.
Ist March,
1932.
9th January,
1933.
8th February,
1933.
1st February,
1st March,
1922.
1923.
No comments yet.
Private notes are available after approval.