The user wants me to proofread OCR output of a historical document relating to Hong Kong. The text appears to be a table of fire brigade personnel with columns: Office, Name, Date of Authority, Annual Salary, Appointment, House or Quarters and Allowances, Leave taken during 1932, Date of First Appointment.
The OCR output is messy with broken lines, misaligned columns, and some garbled text. I need to reconstruct this as a proper Markdown table.
Let me analyze the data:
The header row shows:
The data rows seem to be for Fire Brigade personnel. Let me parse each row.
From the OCR text, I can see entries like:
This is very messy. The OCR seems to have read columns incorrectly. Let me try to reconstruct based on the original table structure.
Looking at the first line: "224 (J 124 )" - this might be a page number or reference.
Then dates: "11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982."
These appear to be "Date of First Appointment" for various personnel.
Then the table data starts.
Let me try to organize by person. The table likely has multiple rows, each for a different person.
From the text, I can identify names:
And positions:
The "Do." likely means "Ditto" - same as above.
Let me try to reconstruct the table row by row.
First, the header mentions "FIRE BRIGADE,—Continued." so this is a continuation.
The columns appear to be:
But the OCR has mixed these up.
Let me look at the first data entry: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931. Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927. Do., Leung Ki, 1st May, Do. 468 Do. 1931. Do., Tsang Yin. 1st January, Do. 468 Do. 1932. Do., Li Yan-kwo. Do. Do. 168 Do."
This seems like multiple people listed under "Engineer Class III". The "Do." under Office means same position.
So likely each person is a separate row.
Let me parse:
Row 1: Office: Engineer Class III, Name: Wong Tsung, Date of Authority: ?, Annual Salary: $516, Appointment: Establishment, Quarters: Quarters, Leave: 1st June 1931?, Date of First Appointment: ?
But the text says "Wong Tsung. Yeung Ling, 1st September, 1925." - maybe Yeung Ling is another person? Or "Yeung Ling" is part of name? "Wong Tsung Yeung Ling"? Unlikely.
Actually, looking at the original OCR: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931."
This is confusing. Perhaps the OCR has merged two rows.
Let me consider the possibility that the table has two sub-columns for each main column? No.
Another approach: The OCR text includes a list of dates at the top: "11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982."
These are 18 dates. There are 18 names I counted? Let's count names: Wong Tsung, Yeung Ling, Chau Chi-ling, Leung Ki, Tsang Yin, Li Yan-kwo, Fan Lam, Lau Chenng, Chong Kan, Ip Fat, Yeung Yeung, Wong Fn, Leung Sai-lo, Leung Kam-sni, Au Củo-ling, Tang Kau, Kwok Nguu, Wong Taeung-keung. That's 18 names. Good.
So the dates at the top are likely the "Date of First Appointment" for each person in order.
Let's match them in order:
But the text also has "1st September, 1925" for Wong Tsung? Actually the text says "Wong Tsung. Yeung Ling, 1st September, 1925." That might be the date of authority for Wong Tsung and Yeung Ling.
Let's look at the structure: The table likely has columns: Office, Name, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
The OCR has scrambled the columns. I need to reconstruct the table as best as possible.
Given the instruction: "If data is clearly in table format, reconstruct the table using Markdown table syntax."
I should create a Markdown table with the appropriate columns and fill in the data for each person.
Let me try to parse each person's data from the text.
The text after the dates list:
"Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment.
$516
Quarters.
1st June,
Do.
469
Do.
1931.
Do.,
Chau Chi-ling.
1st January,
Do.
516
Do.
1927.
Do.,
Leung Ki,
1st May,
Do.
468
Do.
1931.
Do.,
Tsang Yin.
1st January,
Do.
468
Do.
1932.
Do.,
Li Yan-kwo.
Do.
Do.
168
Do.
Engine Driver,
Fan Lam.
1st April,
Do.
384
Do.
1928.
Do.,
Lau Chenng.
Do.
Do.
372
Do.
Do.,
Chong Kan.
1st July,
Do.
384
Do.
1922.
Do.,
Ip Fat.
1st March,
Do.
348
Do.
1930.
Do.
Do
Yeung Yeung.
1st May,
Do.
324
Do.
1932.
Wong Fn.
15th October,
Do.
381
Do.
1932.
Do.,
Leung Sai-lo.
17th June,
Do.
504
Do.
1923.
Do.,
Leung Kam-sni.
1st January,
Do.
396
Do.
1927.
Do.,
Do.,
Do.,
Au Củo-ling.
Tang Kau.
1st January,
Do.
348
Do.
↑
1930.
1st June,
Do.
336
Do.
1931.
Kwok Nguu.
15th July,
Do.
396
Do.
1981.
Do..
Wong Taeung-keung.
23rd November,
Do.
396
Do.
ARTI
1982."
This is a stream of tokens. The "Do." means ditto (same as previous). The columns are not aligned.
Let me assume the column order as per header: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
But the header in OCR: "OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. Date of First Appointment."
So columns:
Now, the first entry: "Engineer Class III" is Office. "Do." means same office for next person. Then "Wong Tsung." Name. "Yeung Ling," - maybe another name? Or "Yeung Ling" is the Date of Authority? But "1st September, 1925." follows. So perhaps "Yeung Ling" is a separate person? But then "Do." before Wong Tsung suggests Wong Tsung is second person with same office.
Let's read: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925."
Maybe it's: Office: Engineer Class III, Name: Wong Tsung, Date of Authority: (blank?), Annual Salary: ?, Appointment: ?, Quarters: ?, Leave: ?, Date of First Appointment: 11th January 1920 (from list).
But the text "Yeung Ling, 1st September, 1925." could be the next person: Name: Yeung Ling, Date of Authority: 1st September 1925.
Then "Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931." This might be for Yeung Ling: Appointment: Establishment, Annual Salary: $516, Quarters: Quarters, Leave: 1st June 1931?, Date of First Appointment: 21st October 1927? But "Do. 469 Do. 1931" is confusing.
Alternatively, the data for each person is spread across multiple lines in OCR.
Given the complexity, perhaps the best is to present the table as reconstructed with the information we have, using the list of names and dates of first appointment from the top, and the other details from the stream.
But the user expects me to proofread and output Markdown. I need to produce a clean table.
Let me try to parse systematically.
I'll write a script in my mind to parse the tokens.
Tokens (split by newline and punctuation):
The "Do." likely repeats the previous non-"Do." value for that column. But columns are interleaved.
This is essentially a fixed-width table that OCR read line by line, but the columns are not separated.
Given the difficulty, perhaps the original document is a standard Hong Kong Civil Service List table. I can try to reconstruct based on typical format.
In Hong Kong Blue Books / Civil Service Lists, the table for Fire Brigade usually has columns: Office, Name, Date of Appointment, Annual Salary, Allowances, etc.
But here we have "Date of Authority" maybe "Date of Appointment to present post".
The list of dates at the top (18 dates) are likely the "Date of First Appointment" for each of the 18 individuals.
The names in order of appearance in the text stream:
That's 18 names. Good.
Now, the offices: First group "Engineer Class III" for first 6? Then "Engine Driver" for next 12? Let's see: After Li Yan-kwo, we have "Engine Driver, Fan Lam." So first 6 are Engineer Class III, next 12 are Engine Driver.
But there are 6 Engineers? Names 1-6: Wong Tsung, Yeung Ling, Chau Chi-ling, Leung Ki, Tsang Yin, Li Yan-kwo. That's 6.
Then Engine Drivers: Fan Lam, Lau Chenng, Chong Kan, Ip Fat, Yeung Yeung, Wong Fn, Leung Sai-lo, Leung Kam-sni, Au Củo-ling, Tang Kau, Kwok Nguu, Wong Taeung-keung. That's 12. Total 18.
Now, for each person, we need to extract: Date of Authority, Annual Salary, Appointment, Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
The Date of First Appointment we have from the top list in order.
Now, the stream after each name seems to contain the other fields.
Let's parse for first person: Wong Tsung.
Text: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931."
But "Do." after Engineer Class III might be for the second person (Yeung Ling) same office. Then "Wong Tsung." is name. Then "Yeung Ling," is next name? Actually the punctuation: "Wong Tsung. Yeung Ling, 1st September, 1925." So maybe Wong Tsung's data ends at "Wong Tsung." and then Yeung Ling starts.
But then what are Wong Tsung's fields? They might be before his name? Unlikely.
Perhaps the table is arranged with columns: Office, Name, Date of Authority, Annual Salary, Appointment, Quarters, Leave, Date of First Appointment.
And the OCR read row by row but the columns are not delimited.
Let's assume the first row (after header) corresponds to the first date (11th Jan 1920) and first name Wong Tsung.
But the text starts with "Engineer Class III, Do., Wong Tsung." So maybe the first row: Office: Engineer Class III, Name: Wong Tsung, Date of Authority: (missing?), Annual Salary: ?, etc.
The "Do." might be in the Office column for the second row.
I think the best approach is to create a table with the columns as per header, and fill in the data as we can infer, using the top dates for "Date of First Appointment", and the stream for other fields.
Given the instruction: "Preserve original content — do not add or remove substantive words. Do not rephrase, rewrite, or change the author's style, tone, or word choice. Only correct what is clearly wrong."
But the OCR output is not the original content; it's a garbled extraction. My job is to proofread the OCR output, i.e., correct the OCR errors and reconstruct the original text flow.
The original text is a table. So I should reconstruct the table as it originally appeared.
I need to use the OCR text to infer the original table cells.
Let me try to simulate the original table layout.
The header row:
| OFFICE | NAME | Date of Authority | Annual Salary | Appointment | House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose | Leave taken during 1932 | Date of First Appointment |
|---|
Then data rows.
From the top, there is a line "224 (J 124 )" which might be a page number.
And "FIRE BRIGADE,—Continued." is a section header.
Then the list of dates: "11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982."
This list might be the "Date of First Appointment" column values for each row, but placed at the top for some reason (maybe the OCR read the column vertically). In the original table, the Date of First Appointment is the last column. The OCR might have read the last column first because of column detection.
But the instruction: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense (left-to-right, top-to-bottom)."
So I should reorder the columns to the correct order.
The OCR output seems to have read the table in a mixed order. The dates list appears first, then the rest of the columns.
Let me check the OCR text: It starts with "! Dato OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. | FIRE BRIGADE,—Continued. Date of First Appointment. | 224 (J 124 ) 11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982. Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931. Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927. Do., Leung Ki, 1st May, Do. 468 Do. 1931. Do., Tsang Yin. 1st January, Do. 468 Do. 1932. Do., Li Yan-kwo. Do. Do. 168 Do. Engine Driver, Fan Lam. 1st April, Do. 384 Do. 1928. Do., Lau Chenng. Do. Do. 372 Do. Do., Chong Kan. 1st July, Do. 384 Do. 1922. Do., Ip Fat. 1st March, Do. 348 Do. 1930. Do. Do Yeung Yeung. 1st May, Do. 324 Do. 1932. Wong Fn. 15th October, Do. 381 Do. 1932. Do., Leung Sai-lo. 17th June, Do. 504 Do. 1923. Do., Leung Kam-sni. 1st January, Do. 396 Do. 1927. Do., Do., Do., Au Củo-ling. Tang Kau. 1st January, Do. 348 Do. ↑ 1930. 1st June, Do. 336 Do. 1931. Kwok Nguu. 15th July, Do. 396 Do. 1981. Do.. Wong Taeung-keung. 23rd November, Do. 396 Do. ARTI 1982."
The header line includes "Date of First Appointment." at the end, but then the dates list appears right after the page number.
It seems the OCR read the "Date of First Appointment" column first (maybe because it's the last column but OCR read columns right-to-left?).
But the dates list has 18 dates that follows the header includes dates that match the "Date of First Appointment" for each person.
Then the rest of the text includes the other columns for each person, but interleaved.
Let me try to map each person's data from the stream.
I'll write a parsing algorithm manually.
Define columns:
We have 18 persons.
We have the Date of First Appointment from the list (in order). Let's list them with corrected years (1982 -> 1932, 1981 -> 1931? but 1981 appears for Kwok Nguu, but the list has 15th July 1931 for 17th person? Actually the list: 15th July, 1931 is the 17th date. And 23rd November, 1982 is 18th date (should be 1932). So the list is likely the Date of First Appointment column.
Now, the stream after the dates list seems to contain the other columns for each person in order.
Let's break the stream into per-person chunks.
The stream starts with "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931. Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927. Do., Leung Ki, 1st May, Do. 468 Do. 1931. Do., Tsang Yin. 1st January, Do. 468 Do. 1932. Do., Li Yan-kwo. Do. Do. 168 Do. Engine Driver, Fan Lam. 1st April, Do. 384 Do. 1928. Do., Lau Chenng. Do. Do. 372 Do. Do., Chong Kan. 1st July, Do. 384 Do. 1922. Do., Ip Fat. 1st March, Do. 348 Do. 1930. Do. Do Yeung Yeung. 1st May, Do. 324 Do. 1932. Wong Fn. 15th October, Do. 381 Do. 1932. Do., Leung Sai-lo. 17th June, Do. 504 Do. 1923. Do., Leung Kam-sni. 1st January, Do. 396 Do. 1927. Do., Do., Do., Au Củo-ling. Tang Kau. 1st January, Do. 348 Do. ↑ 1930. 1st June, Do. 336 Do. 1931. Kwok Nguu. 15th July, Do. 396 Do. 1981. Do.. Wong Taeung-keung. 23rd November, Do. 396 Do. ARTI 1982."
Notice "Do." appears many times. In the original table, "Do." means "ditto" - same as the cell above.
So for each column, if the value is same as previous row, they write "Do.".
Thus, the table likely has many ditto marks.
We need to reconstruct the table with ditto marks expanded or kept as "Do.".
The instruction: "Preserve original content" - so we should keep "Do." as in original.
But the OCR has lost the column alignment. We need to realign.
Let's assume the original table rows are in order of the Date of First Appointment list.
Row 1: Date of First Appointment: 11th January 1920
Row 2: 21st October 1927
Row 3: 2nd January 1924
Row 4: 7th August 1925
Row 5: 23rd May 1925
Row 6: 19th April 1928
Row 7: 6th March 1924
Row 8: 31st January 1922
Row 9: 1st December 1921
Row 10: 1st May 1927
Row 11: 1st January 1925
Row 12: 17th November 1923
Row 13: 17th June 1923
Row 14: 1st April 1926
Row 15: 5th November 1928
Row 16: 15th September 1927
Row 17: 15th July 1931
Row 18: 23rd November 1932
Now, the names in order from stream:
Offices: First 6 are Engineer Class III, next 12 are Engine Driver.
Now, for each row, we need to extract the other columns from the stream.
The stream seems to be a linearized version of the table, perhaps reading row by row but only non-ditto cells? Or reading column by column?
Let's look at the beginning of stream: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931."
This might correspond to the first two rows.
Row1 (Wong Tsung): Office: Engineer Class III, Name: Wong Tsung, Date of Authority: ?, Annual Salary: ?, Appointment: ?, Quarters: ?, Leave: ?, Date of First Appointment: 11 Jan 1920.
Row2 (Yeung Ling): Office: Do. (Engineer Class III), Name: Yeung Ling, Date of Authority: 1st September 1925, Annual Salary: ?, Appointment: Establishment, Salary: $516, Quarters: Quarters, Leave: 1st June 1931?, Date of First Appointment: 21 Oct 1927.
But then "Do. 469 Do. 1931." might be for row2? "Do." for Appointment? "469" for Salary? "Do." for Quarters? "1931" for Leave? But Leave is "1st June" already.
This is messy.
Perhaps the stream is actually the table read column by column? For example, the Office column for all rows, then Name column, then Date of Authority, etc.
But the stream has mixed.
Let's check the header order in OCR: "OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. Date of First Appointment."
The OCR then outputs "Date of First Appointment." and then the dates list. So it read the last column first.
Then it outputs the rest? But the stream starts with "Engineer Class III, Do., Wong Tsung." which are Office and Name for first two rows.
Then "Yeung Ling, 1st September, 1925." - that's Name and Date of Authority for second row? But Yeung Ling is name for row2, Date of Authority 1st Sept 1925.
Then "Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931." - could be Appointment, Annual Salary, Quarters, Leave, ... for row2?
Then "Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927." - Office (Do.), Name (Chau Chi-ling), Date of Authority (1st January), Annual Salary (516), Quarters (Do.), Leave (1927?) but 1927 is a year, maybe Date of First Appointment? But that's already known.
Wait, "1927" appears as a year in the stream for Chau Chi-ling. But Chau Chi-ling's Date of First Appointment is 2nd January 1924 (from list). So 1927 is not that.
Maybe the stream includes the Date of Authority and Annual Salary etc.
Let's try to parse the stream as a sequence of cells in row-major order, but with ditto marks.
We have 18 rows, 8 columns = 144 cells. But many are ditto.
The stream has far fewer tokens.
Maybe the OCR only captured non-ditto cells? But it captured many "Do.".
Another idea: The original table might have been printed with vertical lines, and OCR read it as text lines, but the columns are not separated. The text we see is the raw OCR lines.
Look at the first few lines of OCR input:
"OFFICE.
NAME.
of
Authority.
Aummal Salary.
Appointment,
House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
Leave taken
Juring 1932.
FIRE BRIGADE,—Continued.
Date of First Appointment.
224
(J 124 )
11th January, 1920, 21st October,
31st January,
1922.
1st December, 1921. 1st May, 1927.
1st January, 1925.
17th November.
1st April, 1926.
5th November, !928.
15th Septemer, 1927. 15th July,
1931.
23rd November,
1982.
Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment.
$516
Quarters.
1st June,
Do.
469
Do.
1931.
Do.,
Chau Chi-ling.
1st January,
Do.
516
Do.
1927.
Do.,
Leung Ki,
1st May,
Do.
468
Do.
1931.
Do.,
Tsang Yin.
1st January,
Do.
468
Do.
1932.
Do.,
Li Yan-kwo.
Do.
Do.
168
Do.
Engine Driver,
Fan Lam.
1st April,
Do.
384
Do.
1928.
Do.,
Lau Chenng.
Do.
Do.
372
Do.
Do.,
Chong Kan.
1st July,
Do.
384
Do.
1922.
Do.,
Ip Fat.
1st March,
Do.
348
Do.
1930.
Do.
Do
Yeung Yeung.
1st May,
Do.
324
Do.
1932.
Wong Fn.
15th October,
Do.
381
Do.
1932.
Do.,
Leung Sai-lo.
17th June,
Do.
504
Do.
1923.
Do.,
Leung Kam-sni.
1st January,
Do.
396
Do.
1927.
Do.,
Do.,
Do.,
Au Củo-ling.
Tang Kau.
1st January,
Do.
348
Do.
↑
1930.
1st June,
Do.
336
Do.
1931.
Kwok Nguu.
15th July,
Do.
396
Do.
1981.
Do..
Wong Taeung-keung.
23rd November,
Do.
396
Do.
ARTI
1982."
This looks like the OCR output lines. The vertical bars "|" might indicate table cell separators? But there are only two "|" in the header.
Actually, the header lines are broken: "OFFICE." on one line, "NAME." on next, "of" on next, "Authority." on next, etc. So the header row is split across lines.
Then "FIRE BRIGADE,—Continued." on a line, "Date of First Appointment." on next.
Then "224 (J 124 )" on a line.
Then the dates list, each date on separate lines or combined.
Then the data lines: each cell on a new line? For example, "Engineer Class III," on a line, "Do.," on next, "Wong Tsung." on next, "Yeung Ling," on next, "1st September," on next, "1925." on next, "Establishment." on next, "$516" on next, "Quarters." on next, "1st June," on next, "Do." on next, "469" on next, "Do." on next, "1931." on next, "Do.," on next, "Chau Chi-ling." on next, etc.
So the OCR has output each cell as a separate line! That's great. The lines correspond to cells in reading order (left to right, top to bottom).
But the header is also split into multiple lines per cell? "OFFICE." is one cell, "NAME." is next, "of Authority." might be two cells? Actually "of Authority." might be "Date of Authority" split: "of" and "Authority."? But the header says "Date of Authority." So maybe the cell is "Date of Authority" but OCR split "Date" missing? The header lines: "OFFICE.", "NAME.", "of", "Authority.", "Aummal Salary.", "Appointment,", "House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.", "Leave taken", "Juring 1932.", "|", "FIRE BRIGADE,—Continued.", "Date of First Appointment.", "|", "224", "(J 124 )", then dates.
The "|" might be table borders.
But the data lines after that are each cell content.
Let's assume the table has 9 columns? Header cells: 1. OFFICE, 2. NAME, 3. Date of Authority, 4. Annual Salary, 5. Appointment, 6. House or Quarters..., 7. Leave taken during 1932, 8. Date of First Appointment. That's 8 columns. But there is a "|" before FIRE BRIGADE and before 224, maybe indicating a row header or something.
The data lines: starting from "Engineer Class III," that's first cell of first data row (Office). Then "Do." second row Office? Or second cell of first row? Since each line is a cell, we need to know the number of columns to map lines to rows.
We have 18 rows (from 18 names). 8 columns => 144 cells. But the data lines count: let's count lines from "Engineer Class III," to end.
I'll count roughly:
1 Engineer Class III,
2 Do.,
3 Wong Tsung.
4 Yeung Ling,
5 1st September,
6 1925.
7 Establishment.
8 $516
9 Quarters.
10 1st June,
11 Do.
12 469
13 Do.
14 1931.
15 Do.,
16 Chau Chi-ling.
17 1st January,
18 Do.
19 516
20 Do.
21 1927.
22 Do.,
23 Leung Ki,
24 1st May,
25 Do.
26 468
27 Do.
28 1931.
29 Do.,
30 Tsang Yin.
31 1st January,
32 Do.
33 468
34 Do.
35 1932.
36 Do.,
37 Li Yan-kwo.
38 Do.
39 Do.
40 168
41 Do.
42 Engine Driver,
43 Fan Lam.
44 1st April,
45 Do.
46 384
47 Do.
48 1928.
49 Do.,
50 Lau Chenng.
51 Do.
52 Do.
53 372
54 Do.
55 Do.,
56 Chong Kan.
57 1st July,
58 Do.
59 384
60 Do.
61 1922.
62 Do.,
63 Ip Fat.
64 1st March,
65 Do.
66 348
67 Do.
68 1930.
69 Do.
70 Do
71 Yeung Yeung.
72 1st May,
73 Do.
74 324
75 Do.
76 1932.
77 Wong Fn.
78 15th October,
79 Do.
80 381
81 Do.
82 1932.
83 Do.,
84 Leung Sai-lo.
85 17th June,
86 Do.
87 504
88 Do.
89 1923.
90 Do.,
91 Leung Kam-sni.
92 1st January,
93 Do.
94 396
95 Do.
96 1927.
97 Do.,
98 Do.,
99 Do.,
100 Au Củo-ling.
101 Tang Kau.
102 1st January,
103 Do.
104 348
105 Do.
106 ↑
107 1930.
108 1st June,
109 Do.
110 336
111 Do.
112 1931.
113 Kwok Nguu.
114 15th July,
115 Do.
116 396
117 Do.
118 1981.
119 Do..
120 Wong Taeung-keung.
121 23rd November,
122 Do.
123 396
124 Do.
125 ARTI
126 1982.
That's 126 lines. 126 / 18 = 7. So maybe 7 columns? But header suggests 8. Let's see: 126 cells for 18 rows => 7 columns exactly. 18*7=126. Good.
So the table has 7 columns, not 8. The header might have 7 columns. Let's check header lines:
1 OFFICE.
2 NAME.
3 of
4 Authority.
5 Aummal Salary.
6 Appointment,
7 House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
8 Leave taken
9 Juring 1932.
10 |
11 FIRE BRIGADE,—Continued.
12 Date of First Appointment.
13 |
14 224
15 (J 124 )
That's 15 lines before dates. But the dates list might be part of the table? Actually the dates list might be the "Date of First Appointment" column for each row, but they appear before the data rows. In the line-by-line cell output, the dates list might be the cells of the last column for all rows, but placed before the other columns? That would be column-major order for that column.
But the data lines we counted (126) start after the dates list. The dates list has 18 dates (each date might be multiple lines). Let's count date lines: from "11th January, 1920," to "23rd November, 1982." That's many lines. But they are not in the 126 count.
If the table is 18 rows x 7 columns = 126 cells, and we have 126 data lines after the dates, then the dates list might be separate (maybe the Date of First Appointment column was read first).
But the header includes "Date of First Appointment." as a column. In the line-by-line output, the header cells are also listed line by line. The header lines before the first "|" might be the first row of the table (header row). But they are split across lines.
Let's parse the header as a row of cells. The lines before the first "|" (which appears after "Juring 1932.") are:
1 OFFICE.
2 NAME.
3 of
4 Authority.
5 Aummal Salary.
6 Appointment,
7 House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
8 Leave taken
9 Juring 1932.
That's 9 lines. But we expect 7 columns. Maybe "of" and "Authority." are two cells? But "Date of Authority" is one column. Could be "Date" missing? The OCR might have missed "Date". The line "of" might be "Date of" but "Date" not captured. "Authority." is separate.
Similarly, "Leave taken" and "Juring 1932." might be two cells? But "Leave taken during 1932" is one column.
The "|" might indicate a new row or a separator.
After the first "|", we have "FIRE BRIGADE,—Continued." and "Date of First Appointment." and another "|". That might be a second header row? Or a title row.
Then "224" and "(J 124 )" might be page number.
Then the dates list: 18 dates, each date might be a cell for the "Date of First Appointment" column for each row. But they appear before the data rows.
Then the data rows follow, each row having 7 cells (since 126 lines / 18 = 7). But the data lines we listed start with "Engineer Class III," which would be the first cell of first data row (Office). Then "Do." second cell? But second cell should be Name. But "Do." is not a name. So maybe the data rows have 7 cells but the columns are different.
Let's map the 7 columns based on the data.
We have 18 rows. For each row, we have 7 cells in order.
Row1 cells (lines 1-7):
Row2 cells (lines 8-14):
Row3 cells (lines 15-21):
Row4 cells (lines 22-28):
Row5 cells (lines 29-35):
Row6 cells (lines 36-42):
Row7 cells (lines 43-49):
Row8 cells (lines 50-56):
Row9 cells (lines 57-63):
Row10 cells (lines 64-70):
Row11 cells (lines 71-77):
Row12 cells (lines 78-84):
Row13 cells (lines 85-91):
Row14 cells (lines 92-98):
Row15 cells (lines 99-105):
Row16 cells (lines 106-112):
Row17 cells (lines 113-119):
Row18 cells (lines 120-126):
This mapping gives 7 cells per row. But the content seems misaligned: Row1 has "Engineer Class III", "Do.", "Wong Tsung", "Yeung Ling", "1st September", "1925", "Establishment". That's 7 cells. But "Do." in second cell? That would be the second column. If columns are: 1. Office, 2. Name, 3. Date of Authority, 4. Annual Salary, 5. Appointment, 6. Quarters, 7. Leave? But then row1 has two names? "Wong Tsung" and "Yeung Ling". And "Do." in column2 is weird.
Maybe the table has a multi-row header or the first row is not a data row but a sub-header.
Notice "Engineer Class III" might be a section header, not a data row. In the original table, there might be a row for "Engineer Class III" as a grade, then the names under it.
But the OCR reads line by line, so "Engineer Class III" is a cell in the first column of the first data row. Then "Do." in the second column? That doesn't make sense.
Perhaps the table has 8 columns, and the first column is "Office", second "Name", third "Date of Authority", fourth "Annual Salary", fifth "Appointment", sixth "House/Quarters", seventh "Leave", eighth "Date of First Appointment". But we have 7 cells per row in the data lines. However, the dates list (18 dates) might be the eighth column, read separately.
If we combine the dates list as the eighth column for each row, then we have 8 columns.
The dates list in order:
Now, the data rows (7 cells each) plus this date as 8th column.
But the data rows as segmented above don't align well with names.
Let's look at the names we expect: 18 names. In the segmented rows, names appear in various cells.
Row1 cell3: Wong Tsung
Row1 cell4: Yeung Ling
Row2: no name
Row3 cell2: Chau Chi-ling
Row4 cell2: Leung Ki
Row5 cell2: Tsang Yin
Row6 cell2: Li Yan-kwo
Row7 cell1: Fan Lam
Row8 cell1: Lau Chenng
Row9 cell6: Ip Fat? Actually row9 cell6 is "Do.," cell7 is "Ip Fat." Wait row9 cells: 57. 1st July, 58. Do. 59. 384 60. Do. 61. 1922. 62. Do., 63. Ip Fat. So cell7 is Ip Fat.
Row10 cell6: "Do" (line70) and cell7? Row10 has 7 cells: lines64-70: 64. 1st March, 65. Do. 66. 348 67. Do. 68. 1930. 69. Do. 70. Do. So no name in row10.
Row11 cell1: Yeung Yeung.
Row11 cell7: Wong Fn.
Row12 cell7: Leung Sai-lo.
Row13 cell7: Leung Kam-sni.
Row14 cell7: Do. (line98)
Row15 cell2: Au Củo-ling. cell3: Tang Kau.
Row16: no clear name.
Row17 cell1: Kwok Nguu.
Row18 cell1: Wong Taeung-keung.
This is messy. The segmentation into 7-cell rows might be wrong because the OCR lines might not correspond exactly to cells; some cells might span multiple lines (like "House or Quarters..." is one cell but multiple lines in header). In data, a cell might be multiple lines? But the data lines seem each line is a cell.
Maybe the table has 9 columns? 126/18=7 exactly, so 7 columns.
Let's examine the header again. The header lines before the first "|" are 9 lines. But two of them might be continuations: "of" and "Authority." belong together as "Date of Authority". "Leave taken" and "Juring 1932." belong together as "Leave taken during 1932". So that reduces to 7 header cells:
That's 7 columns. The "Date of First Appointment" appears after the "|" as a separate header? But there is a "|" then "FIRE BRIGADE,—Continued." then "Date of First Appointment." then "|". That might be a title row and a column header for the last column? But we have only 7 columns in data.
Maybe the "Date of First Appointment" is not a column in this table? But the header includes it. The OCR shows "Date of First Appointment." after the second "|". The first "|" might be a vertical line in the table separating the main columns from a remarks column? Or the table has two sections.
Look at the original OCR: "Leave taken Juring 1932. | FIRE BRIGADE,—Continued. Date of First Appointment. | 224 (J 124 )". The "|" might be table borders. So the table has columns up to "Leave taken during 1932", then a border, then "FIRE BRIGADE,—Continued." might be a row header? And "Date of First Appointment." might be another column? But then "224 (J 124 )" might be page number.
Alternatively, the table might be continued from previous page, and "FIRE BRIGADE,—Continued." is a center heading. The "Date of First Appointment" might be the last column of the table.
Given the data lines have 7 cells per row, and the header has 7 meaningful columns (if we combine split lines), I'll assume the table has 7 columns: Office, Name, Date of Authority, Annual Salary, Appointment, Quarters/Allowances, Leave taken during 1932. The Date of First Appointment might be in a separate column but not captured in the data lines? But the dates list is there.
Wait, the dates list has 18 entries. If the table has 18 rows, and Date of First Appointment is a column, then the data lines should have 8 cells per row. But we have 126 lines, 126/18=7. So either the Date of First Appointment is not in the data lines (it's in the dates list), or the data lines include it but we miscounted.
Let's count the data lines again carefully. The data lines start after "23rd November, 1982." The next line is "Engineer Class III,". Let's number from there to end.
I'll copy the data lines as a list:
Yes, 126 lines.
Now, if the table has 7 columns, then each row has 7 cells. 18 rows * 7 = 126. So the mapping of lines to rows is sequential: lines 1-7 row1, 8-14 row2, etc.
But the content of row1: cells:
Row2:
Row3:
Row4:
Row5:
Row6:
Row7:
Row8:
Row9:
Row10:
Row11:
Row12:
Row13:
Row14:
Row15:
Row16:
Row17:
Row18:
Now, we need to assign column headers to these 7 columns.
From the header lines (combined):
Col1: OFFICE
Col2: NAME
Col3: Date of Authority
Col4: Annual Salary
Col5: Appointment
Col6: House or Quarters, and Allowances...
Col7: Leave taken during 1932
But the data in row1 col1 is "Engineer Class III" (an office), col2 is "Do." (ditto for office? but col2 is Name). That doesn't match.
Maybe the columns are different. Let's look at row1: col1: Engineer Class III (office), col2: Do. (maybe ditto for office? but then col3: Wong Tsung (name), col4: Yeung Ling (name?), col5: 1st September (date), col6: 1925 (year), col7: Establishment (appointment). That suggests the table might have two name columns? Or the first row is a header for a group.
In many civil service lists, there is a row for the grade "Engineer Class III" and then the officers listed below with their names. But the table structure might have the grade in the first column, and the name in the second column, but the grade row might have the grade in first column and the rest blank or "Do.".
But here row1 has "Engineer Class III" in col1, "Do." in col2, "Wong Tsung" in col3, "Yeung Ling" in col4. That's odd.
Perhaps the OCR has misordered the columns. The instruction says: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense."
The OCR read the table line by line, but the table might have been arranged with columns not strictly left-to-right in the OCR output. However, the OCR output lines seem to be in reading order (left to right, top to bottom) for each cell. But the table might have a complex layout with merged cells.
Given the difficulty, perhaps the best is to reconstruct the table as it appears in the original by using the known structure of Hong Kong Civil Service Lists.
I recall that the Hong Kong Blue Book / Civil Service List has tables with columns: Office, Name, Date of Appointment, Annual Salary, Allowances, etc.
But here we have a specific snippet.
Given the user expects a proofread text in Markdown, and the input is OCR output, I should output a cleaned-up version of the text, preserving the table structure.
Since the OCR output is essentially a list of cells in order, I can try to reconstruct the table by assuming the first few lines are header, then the dates list is the "Date of First Appointment" column, then the data lines are the other columns.
But the data lines have 7 cells per row, and we have 18 rows. The dates list has 18 dates. So if we take the dates list as column 8, we have 8 columns.
But the header has 7 columns before the "|". The "Date of First Appointment" appears after the "|". So maybe the table has 8 columns: the first 7 as per header, and the 8th is "Date of First Appointment". The OCR read the 8th column first (the dates list), then the other 7 columns for each row.
That would make sense: The OCR read the table column by column? But the data lines are 126, which is 18*7. So the data lines are the first 7 columns for all rows, in row-major order. The dates list is the 8th column, read separately (maybe because it was at the rightmost and OCR read it first).
Thus, the original table has 8 columns. The header row has 8 cells: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
The OCR output: first, the header cells for the first 7 columns (split across lines), then a "|" (border), then "FIRE BRIGADE,—Continued." (maybe a title), then "Date of First Appointment." (the 8th header), then "|", then page number, then the 8th column data (18 dates), then the first 7 columns data for all rows (126 cells).
That matches!
So we have:
But the 126 lines are in row-major order for the first 7 columns. However, the first 7 columns of row1 are lines 1-7, row2 lines 8-14, etc.
But row1 col1 is "Engineer Class III", col2 "Do.", col3 "Wong Tsung", col4 "Yeung Ling", col5 "1st September", col6 "1925", col7 "Establishment". That would be the first 7 columns for the first row. But the first row should correspond to the first date (11th Jan 1920). The name should be Wong Tsung? But col3 is Wong Tsung, col4 is Yeung Ling. That suggests that the first row actually contains two names? Or the table has two rows for the first two persons in the first row? That doesn't make sense.
Maybe the data lines are not row-major for the first 7 columns, but column-major? Let's check: If the 126 lines are column-major for the first 7 columns, then the first 18 lines would be column1 for all rows, next 18 lines column2, etc. 7 columns * 18 rows = 126. So lines 1-18: column1 (Office), lines 19-36: column2 (Name), lines 37-54: column3 (Date of Authority), lines 55-72: column4 (Annual Salary), lines 73-90: column5 (Appointment), lines 91-108: column6 (Quarters), lines 109-126: column7 (Leave).
Let's test this hypothesis.
Lines 1-18 (Office column):
That doesn't look like all offices. Line3 is "Wong Tsung." which is a name, not office. So not column-major.
Maybe the data lines are in the order of the original table reading order (left to right, top to bottom) but the table has a different number of columns per row? Or the table has merged cells.
Given the time, I think the most practical is to output the text as a cleaned-up version of the OCR, preserving the table structure as best as possible, using Markdown table with the columns as per header, and filling in the data from the OCR text in a logical way.
But the user said: "Your most important task is to fix the text flow while preserving the original paragraph structure." This is a table, not paragraphs. So I should reconstruct the table.
I will create a Markdown table with the 8 columns. I'll use the dates list for the last column. For the first 7 columns, I'll try to parse the 126 lines into 18 rows of 7 columns by assuming the OCR read the table row by row, but the first row might be a header row for the group "Engineer Class III". Actually, the table might have a row for "Engineer Class III" spanning multiple columns? But the OCR reads cells linearly.
Let's look at the original OCR text again: It starts with "OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. | FIRE BRIGADE,—Continued. Date of First Appointment. | 224 (J 124 )". This suggests the table has a header row with those columns, then a title "FIRE BRIGADE,—Continued." maybe centered, then "Date of First Appointment" might be a column header for the last column? But it's after a "|".
Maybe the table is split into two parts: left part with columns OFFICE to Leave taken, right part with Date of First Appointment. The "|" indicates a vertical line.
In the data, the dates list appears after the page number, then the left part data.
The left part data (126 lines) might be the left part columns for all rows. The left part has 7 columns (OFFICE to Leave taken). 18 rows * 7 = 126. Good.
The right part (Date of First Appointment) has 18 dates.
So the full table has 8 columns. The left 7 columns data is in the 126 lines in row-major order. The rightmost column data is in the dates list in order.
Now, we need to split the 126 lines into 18 rows of 7 columns each. We already did that. But the content of row1 (lines1-7) seems to contain data for two persons? Let's see row1:
Col1: Engineer Class III
Col2: Do.
Col3: Wong Tsung
Col4: Yeung Ling
Col5: 1st September
Col6: 1925
Col7: Establishment
But the left part columns are: 1. OFFICE, 2. NAME, 3. Date of Authority, 4. Annual Salary, 5. Appointment, 6. House/Quarters/Allowances, 7. Leave taken during 1932.
So for row1, OFFICE = "Engineer Class III", NAME = "Do."? That can't be. NAME should be a name. "Do." means ditto, so same as previous row's name. But there is no previous row. So maybe the first row is not a data row but a sub-heading. In many tables, the grade "Engineer Class III" is written in the OFFICE column for the first officer, and then "Do." for subsequent officers. But here the first officer is Wong Tsung. So row1 should be: OFFICE: Engineer Class III, NAME: Wong Tsung, Date of Authority: ?, Annual Salary: ?, Appointment: ?, Quarters: ?, Leave: ?.
But in our segmentation, row1 has NAME = "Do." and then Wong Tsung appears in col3. So the columns are shifted.
Perhaps the table has an extra column? Or the "Do." in col2 is actually the Date of Authority for the previous? No.
Let's look at the raw lines again. The first few lines after the dates list:
"Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment."
If we consider that the table might have been read with a different column order: maybe the OCR read the table in a different order because of the
The user wants me to proofread OCR output of a historical document relating to Hong Kong. The text appears to be a table of fire brigade personnel with columns: Office, Name, Date of Authority, Annual Salary, Appointment, House or Quarters and Allowances, Leave taken during 1932, Date of First Appointment.
The OCR output is messy with broken lines, misaligned columns, and some garbled text. I need to reconstruct this as a proper Markdown table.
Let me analyze the data:
The header row shows:
The data rows seem to be for Fire Brigade personnel. Let me parse each row.
From the OCR text, I can see entries like:
This is very messy. The OCR seems to have read columns incorrectly. Let me try to reconstruct based on the original table structure.
Looking at the first line: "224 (J 124 )" - this might be a page number or reference.
Then dates: "11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982."
These appear to be "Date of First Appointment" for various personnel.
Then the table data starts.
Let me try to organize by person. The table likely has multiple rows, each for a different person.
From the text, I can identify names:
And positions:
The "Do." likely means "Ditto" - same as above.
Let me try to reconstruct the table row by row.
First, the header mentions "FIRE BRIGADE,—Continued." so this is a continuation.
The columns appear to be:
But the OCR has mixed these up.
Let me look at the first data entry: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931. Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927. Do., Leung Ki, 1st May, Do. 468 Do. 1931. Do., Tsang Yin. 1st January, Do. 468 Do. 1932. Do., Li Yan-kwo. Do. Do. 168 Do."
This seems like multiple people listed under "Engineer Class III". The "Do." under Office means same position.
So likely each person is a separate row.
Let me parse:
Row 1: Office: Engineer Class III, Name: Wong Tsung, Date of Authority: ?, Annual Salary: $516, Appointment: Establishment, Quarters: Quarters, Leave: 1st June 1931?, Date of First Appointment: ?
But the text says "Wong Tsung. Yeung Ling, 1st September, 1925." - maybe Yeung Ling is another person? Or "Yeung Ling" is part of name? "Wong Tsung Yeung Ling"? Unlikely.
Actually, looking at the original OCR: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931."
This is confusing. Perhaps the OCR has merged two rows.
Let me consider the possibility that the table has two sub-columns for each main column? No.
Another approach: The OCR text includes a list of dates at the top: "11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982."
These are 18 dates. There are 18 names I counted? Let's count names: Wong Tsung, Yeung Ling, Chau Chi-ling, Leung Ki, Tsang Yin, Li Yan-kwo, Fan Lam, Lau Chenng, Chong Kan, Ip Fat, Yeung Yeung, Wong Fn, Leung Sai-lo, Leung Kam-sni, Au Củo-ling, Tang Kau, Kwok Nguu, Wong Taeung-keung. That's 18 names. Good.
So the dates at the top are likely the "Date of First Appointment" for each person in order.
Let's match them in order:
But the text also has "1st September, 1925" for Wong Tsung? Actually the text says "Wong Tsung. Yeung Ling, 1st September, 1925." That might be the date of authority for Wong Tsung and Yeung Ling.
Let's look at the structure: The table likely has columns: Office, Name, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
The OCR has scrambled the columns. I need to reconstruct the table as best as possible.
Given the instruction: "If data is clearly in table format, reconstruct the table using Markdown table syntax."
I should create a Markdown table with the appropriate columns and fill in the data for each person.
Let me try to parse each person's data from the text.
The text after the dates list:
"Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment.
$516
Quarters.
1st June,
Do.
469
Do.
1931.
Do.,
Chau Chi-ling.
1st January,
Do.
516
Do.
1927.
Do.,
Leung Ki,
1st May,
Do.
468
Do.
1931.
Do.,
Tsang Yin.
1st January,
Do.
468
Do.
1932.
Do.,
Li Yan-kwo.
Do.
Do.
168
Do.
Engine Driver,
Fan Lam.
1st April,
Do.
384
Do.
1928.
Do.,
Lau Chenng.
Do.
Do.
372
Do.
Do.,
Chong Kan.
1st July,
Do.
384
Do.
1922.
Do.,
Ip Fat.
1st March,
Do.
348
Do.
1930.
Do.
Do
Yeung Yeung.
1st May,
Do.
324
Do.
1932.
Wong Fn.
15th October,
Do.
381
Do.
1932.
Do.,
Leung Sai-lo.
17th June,
Do.
504
Do.
1923.
Do.,
Leung Kam-sni.
1st January,
Do.
396
Do.
1927.
Do.,
Do.,
Do.,
Au Củo-ling.
Tang Kau.
1st January,
Do.
348
Do.
↑
1930.
1st June,
Do.
336
Do.
1931.
Kwok Nguu.
15th July,
Do.
396
Do.
1981.
Do..
Wong Taeung-keung.
23rd November,
Do.
396
Do.
ARTI
1982."
This is a stream of tokens. The "Do." means ditto (same as previous). The columns are not aligned.
Let me assume the column order as per header: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
But the header in OCR: "OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. Date of First Appointment."
So columns:
Now, the first entry: "Engineer Class III" is Office. "Do." means same office for next person. Then "Wong Tsung." Name. "Yeung Ling," - maybe another name? Or "Yeung Ling" is the Date of Authority? But "1st September, 1925." follows. So perhaps "Yeung Ling" is a separate person? But then "Do." before Wong Tsung suggests Wong Tsung is second person with same office.
Let's read: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925."
Maybe it's: Office: Engineer Class III, Name: Wong Tsung, Date of Authority: (blank?), Annual Salary: ?, Appointment: ?, Quarters: ?, Leave: ?, Date of First Appointment: 11th January 1920 (from list).
But the text "Yeung Ling, 1st September, 1925." could be the next person: Name: Yeung Ling, Date of Authority: 1st September 1925.
Then "Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931." This might be for Yeung Ling: Appointment: Establishment, Annual Salary: $516, Quarters: Quarters, Leave: 1st June 1931?, Date of First Appointment: 21st October 1927? But "Do. 469 Do. 1931" is confusing.
Alternatively, the data for each person is spread across multiple lines in OCR.
Given the complexity, perhaps the best is to present the table as reconstructed with the information we have, using the list of names and dates of first appointment from the top, and the other details from the stream.
But the user expects me to proofread and output Markdown. I need to produce a clean table.
Let me try to parse systematically.
I'll write a script in my mind to parse the tokens.
Tokens (split by newline and punctuation):
The "Do." likely repeats the previous non-"Do." value for that column. But columns are interleaved.
This is essentially a fixed-width table that OCR read line by line, but the columns are not separated.
Given the difficulty, perhaps the original document is a standard Hong Kong Civil Service List table. I can try to reconstruct based on typical format.
In Hong Kong Blue Books / Civil Service Lists, the table for Fire Brigade usually has columns: Office, Name, Date of Appointment, Annual Salary, Allowances, etc.
But here we have "Date of Authority" maybe "Date of Appointment to present post".
The list of dates at the top (18 dates) are likely the "Date of First Appointment" for each of the 18 individuals.
The names in order of appearance in the text stream:
That's 18 names. Good.
Now, the offices: First group "Engineer Class III" for first 6? Then "Engine Driver" for next 12? Let's see: After Li Yan-kwo, we have "Engine Driver, Fan Lam." So first 6 are Engineer Class III, next 12 are Engine Driver.
But there are 6 Engineers? Names 1-6: Wong Tsung, Yeung Ling, Chau Chi-ling, Leung Ki, Tsang Yin, Li Yan-kwo. That's 6.
Then Engine Drivers: Fan Lam, Lau Chenng, Chong Kan, Ip Fat, Yeung Yeung, Wong Fn, Leung Sai-lo, Leung Kam-sni, Au Củo-ling, Tang Kau, Kwok Nguu, Wong Taeung-keung. That's 12. Total 18.
Now, for each person, we need to extract: Date of Authority, Annual Salary, Appointment, Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
The Date of First Appointment we have from the top list in order.
Now, the stream after each name seems to contain the other fields.
Let's parse for first person: Wong Tsung.
Text: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931."
But "Do." after Engineer Class III might be for the second person (Yeung Ling) same office. Then "Wong Tsung." is name. Then "Yeung Ling," is next name? Actually the punctuation: "Wong Tsung. Yeung Ling, 1st September, 1925." So maybe Wong Tsung's data ends at "Wong Tsung." and then Yeung Ling starts.
But then what are Wong Tsung's fields? They might be before his name? Unlikely.
Perhaps the table is arranged with columns: Office, Name, Date of Authority, Annual Salary, Appointment, Quarters, Leave, Date of First Appointment.
And the OCR read row by row but the columns are not delimited.
Let's assume the first row (after header) corresponds to the first date (11th Jan 1920) and first name Wong Tsung.
But the text starts with "Engineer Class III, Do., Wong Tsung." So maybe the first row: Office: Engineer Class III, Name: Wong Tsung, Date of Authority: (missing?), Annual Salary: ?, etc.
The "Do." might be in the Office column for the second row.
I think the best approach is to create a table with the columns as per header, and fill in the data as we can infer, using the top dates for "Date of First Appointment", and the stream for other fields.
Given the instruction: "Preserve original content — do not add or remove substantive words. Do not rephrase, rewrite, or change the author's style, tone, or word choice. Only correct what is clearly wrong."
But the OCR output is not the original content; it's a garbled extraction. My job is to proofread the OCR output, i.e., correct the OCR errors and reconstruct the original text flow.
The original text is a table. So I should reconstruct the table as it originally appeared.
I need to use the OCR text to infer the original table cells.
Let me try to simulate the original table layout.
The header row:
| OFFICE | NAME | Date of Authority | Annual Salary | Appointment | House or Quarters, and Allowances for Rent, Entertainment, Personal, or for any other purpose | Leave taken during 1932 | Date of First Appointment |
|---|
Then data rows.
From the top, there is a line "224 (J 124 )" which might be a page number.
And "FIRE BRIGADE,—Continued." is a section header.
Then the list of dates: "11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982."
This list might be the "Date of First Appointment" column values for each row, but placed at the top for some reason (maybe the OCR read the column vertically). In the original table, the Date of First Appointment is the last column. The OCR might have read the last column first because of column detection.
But the instruction: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense (left-to-right, top-to-bottom)."
So I should reorder the columns to the correct order.
The OCR output seems to have read the table in a mixed order. The dates list appears first, then the rest of the columns.
Let me check the OCR text: It starts with "! Dato OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. | FIRE BRIGADE,—Continued. Date of First Appointment. | 224 (J 124 ) 11th January, 1920, 21st October, 1927. 2nd January, 1924. 7th August, 1925. 23rd May, 1925. 19th April, 1928. 6th March, 1924. 31st January, 1922. 1st December, 1921. 1st May, 1927. 1st January, 1925. 17th November. 1923. 17th June, 1923. 1st April, 1926. 5th November, !928. 15th Septemer, 1927. 15th July, 1931. 23rd November, 1982. Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931. Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927. Do., Leung Ki, 1st May, Do. 468 Do. 1931. Do., Tsang Yin. 1st January, Do. 468 Do. 1932. Do., Li Yan-kwo. Do. Do. 168 Do. Engine Driver, Fan Lam. 1st April, Do. 384 Do. 1928. Do., Lau Chenng. Do. Do. 372 Do. Do., Chong Kan. 1st July, Do. 384 Do. 1922. Do., Ip Fat. 1st March, Do. 348 Do. 1930. Do. Do Yeung Yeung. 1st May, Do. 324 Do. 1932. Wong Fn. 15th October, Do. 381 Do. 1932. Do., Leung Sai-lo. 17th June, Do. 504 Do. 1923. Do., Leung Kam-sni. 1st January, Do. 396 Do. 1927. Do., Do., Do., Au Củo-ling. Tang Kau. 1st January, Do. 348 Do. ↑ 1930. 1st June, Do. 336 Do. 1931. Kwok Nguu. 15th July, Do. 396 Do. 1981. Do.. Wong Taeung-keung. 23rd November, Do. 396 Do. ARTI 1982."
The header line includes "Date of First Appointment." at the end, but then the dates list appears right after the page number.
It seems the OCR read the "Date of First Appointment" column first (maybe because it's the last column but OCR read columns right-to-left?).
But the dates list has 18 dates that follows the header includes dates that match the "Date of First Appointment" for each person.
Then the rest of the text includes the other columns for each person, but interleaved.
Let me try to map each person's data from the stream.
I'll write a parsing algorithm manually.
Define columns:
We have 18 persons.
We have the Date of First Appointment from the list (in order). Let's list them with corrected years (1982 -> 1932, 1981 -> 1931? but 1981 appears for Kwok Nguu, but the list has 15th July 1931 for 17th person? Actually the list: 15th July, 1931 is the 17th date. And 23rd November, 1982 is 18th date (should be 1932). So the list is likely the Date of First Appointment column.
Now, the stream after the dates list seems to contain the other columns for each person in order.
Let's break the stream into per-person chunks.
The stream starts with "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931. Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927. Do., Leung Ki, 1st May, Do. 468 Do. 1931. Do., Tsang Yin. 1st January, Do. 468 Do. 1932. Do., Li Yan-kwo. Do. Do. 168 Do. Engine Driver, Fan Lam. 1st April, Do. 384 Do. 1928. Do., Lau Chenng. Do. Do. 372 Do. Do., Chong Kan. 1st July, Do. 384 Do. 1922. Do., Ip Fat. 1st March, Do. 348 Do. 1930. Do. Do Yeung Yeung. 1st May, Do. 324 Do. 1932. Wong Fn. 15th October, Do. 381 Do. 1932. Do., Leung Sai-lo. 17th June, Do. 504 Do. 1923. Do., Leung Kam-sni. 1st January, Do. 396 Do. 1927. Do., Do., Do., Au Củo-ling. Tang Kau. 1st January, Do. 348 Do. ↑ 1930. 1st June, Do. 336 Do. 1931. Kwok Nguu. 15th July, Do. 396 Do. 1981. Do.. Wong Taeung-keung. 23rd November, Do. 396 Do. ARTI 1982."
Notice "Do." appears many times. In the original table, "Do." means "ditto" - same as the cell above.
So for each column, if the value is same as previous row, they write "Do.".
Thus, the table likely has many ditto marks.
We need to reconstruct the table with ditto marks expanded or kept as "Do.".
The instruction: "Preserve original content" - so we should keep "Do." as in original.
But the OCR has lost the column alignment. We need to realign.
Let's assume the original table rows are in order of the Date of First Appointment list.
Row 1: Date of First Appointment: 11th January 1920
Row 2: 21st October 1927
Row 3: 2nd January 1924
Row 4: 7th August 1925
Row 5: 23rd May 1925
Row 6: 19th April 1928
Row 7: 6th March 1924
Row 8: 31st January 1922
Row 9: 1st December 1921
Row 10: 1st May 1927
Row 11: 1st January 1925
Row 12: 17th November 1923
Row 13: 17th June 1923
Row 14: 1st April 1926
Row 15: 5th November 1928
Row 16: 15th September 1927
Row 17: 15th July 1931
Row 18: 23rd November 1932
Now, the names in order from stream:
Offices: First 6 are Engineer Class III, next 12 are Engine Driver.
Now, for each row, we need to extract the other columns from the stream.
The stream seems to be a linearized version of the table, perhaps reading row by row but only non-ditto cells? Or reading column by column?
Let's look at the beginning of stream: "Engineer Class III, Do., Wong Tsung. Yeung Ling, 1st September, 1925. Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931."
This might correspond to the first two rows.
Row1 (Wong Tsung): Office: Engineer Class III, Name: Wong Tsung, Date of Authority: ?, Annual Salary: ?, Appointment: ?, Quarters: ?, Leave: ?, Date of First Appointment: 11 Jan 1920.
Row2 (Yeung Ling): Office: Do. (Engineer Class III), Name: Yeung Ling, Date of Authority: 1st September 1925, Annual Salary: ?, Appointment: Establishment, Salary: $516, Quarters: Quarters, Leave: 1st June 1931?, Date of First Appointment: 21 Oct 1927.
But then "Do. 469 Do. 1931." might be for row2? "Do." for Appointment? "469" for Salary? "Do." for Quarters? "1931" for Leave? But Leave is "1st June" already.
This is messy.
Perhaps the stream is actually the table read column by column? For example, the Office column for all rows, then Name column, then Date of Authority, etc.
But the stream has mixed.
Let's check the header order in OCR: "OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. Date of First Appointment."
The OCR then outputs "Date of First Appointment." and then the dates list. So it read the last column first.
Then it outputs the rest? But the stream starts with "Engineer Class III, Do., Wong Tsung." which are Office and Name for first two rows.
Then "Yeung Ling, 1st September, 1925." - that's Name and Date of Authority for second row? But Yeung Ling is name for row2, Date of Authority 1st Sept 1925.
Then "Establishment. $516 Quarters. 1st June, Do. 469 Do. 1931." - could be Appointment, Annual Salary, Quarters, Leave, ... for row2?
Then "Do., Chau Chi-ling. 1st January, Do. 516 Do. 1927." - Office (Do.), Name (Chau Chi-ling), Date of Authority (1st January), Annual Salary (516), Quarters (Do.), Leave (1927?) but 1927 is a year, maybe Date of First Appointment? But that's already known.
Wait, "1927" appears as a year in the stream for Chau Chi-ling. But Chau Chi-ling's Date of First Appointment is 2nd January 1924 (from list). So 1927 is not that.
Maybe the stream includes the Date of Authority and Annual Salary etc.
Let's try to parse the stream as a sequence of cells in row-major order, but with ditto marks.
We have 18 rows, 8 columns = 144 cells. But many are ditto.
The stream has far fewer tokens.
Maybe the OCR only captured non-ditto cells? But it captured many "Do.".
Another idea: The original table might have been printed with vertical lines, and OCR read it as text lines, but the columns are not separated. The text we see is the raw OCR lines.
Look at the first few lines of OCR input:
"OFFICE.
NAME.
of
Authority.
Aummal Salary.
Appointment,
House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
Leave taken
Juring 1932.
FIRE BRIGADE,—Continued.
Date of First Appointment.
224
(J 124 )
11th January, 1920, 21st October,
31st January,
1922.
1st December, 1921. 1st May, 1927.
1st January, 1925.
17th November.
1st April, 1926.
5th November, !928.
15th Septemer, 1927. 15th July,
1931.
23rd November,
1982.
Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment.
$516
Quarters.
1st June,
Do.
469
Do.
1931.
Do.,
Chau Chi-ling.
1st January,
Do.
516
Do.
1927.
Do.,
Leung Ki,
1st May,
Do.
468
Do.
1931.
Do.,
Tsang Yin.
1st January,
Do.
468
Do.
1932.
Do.,
Li Yan-kwo.
Do.
Do.
168
Do.
Engine Driver,
Fan Lam.
1st April,
Do.
384
Do.
1928.
Do.,
Lau Chenng.
Do.
Do.
372
Do.
Do.,
Chong Kan.
1st July,
Do.
384
Do.
1922.
Do.,
Ip Fat.
1st March,
Do.
348
Do.
1930.
Do.
Do
Yeung Yeung.
1st May,
Do.
324
Do.
1932.
Wong Fn.
15th October,
Do.
381
Do.
1932.
Do.,
Leung Sai-lo.
17th June,
Do.
504
Do.
1923.
Do.,
Leung Kam-sni.
1st January,
Do.
396
Do.
1927.
Do.,
Do.,
Do.,
Au Củo-ling.
Tang Kau.
1st January,
Do.
348
Do.
↑
1930.
1st June,
Do.
336
Do.
1931.
Kwok Nguu.
15th July,
Do.
396
Do.
1981.
Do..
Wong Taeung-keung.
23rd November,
Do.
396
Do.
ARTI
1982."
This looks like the OCR output lines. The vertical bars "|" might indicate table cell separators? But there are only two "|" in the header.
Actually, the header lines are broken: "OFFICE." on one line, "NAME." on next, "of" on next, "Authority." on next, etc. So the header row is split across lines.
Then "FIRE BRIGADE,—Continued." on a line, "Date of First Appointment." on next.
Then "224 (J 124 )" on a line.
Then the dates list, each date on separate lines or combined.
Then the data lines: each cell on a new line? For example, "Engineer Class III," on a line, "Do.," on next, "Wong Tsung." on next, "Yeung Ling," on next, "1st September," on next, "1925." on next, "Establishment." on next, "$516" on next, "Quarters." on next, "1st June," on next, "Do." on next, "469" on next, "Do." on next, "1931." on next, "Do.," on next, "Chau Chi-ling." on next, etc.
So the OCR has output each cell as a separate line! That's great. The lines correspond to cells in reading order (left to right, top to bottom).
But the header is also split into multiple lines per cell? "OFFICE." is one cell, "NAME." is next, "of Authority." might be two cells? Actually "of Authority." might be "Date of Authority" split: "of" and "Authority."? But the header says "Date of Authority." So maybe the cell is "Date of Authority" but OCR split "Date" missing? The header lines: "OFFICE.", "NAME.", "of", "Authority.", "Aummal Salary.", "Appointment,", "House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.", "Leave taken", "Juring 1932.", "|", "FIRE BRIGADE,—Continued.", "Date of First Appointment.", "|", "224", "(J 124 )", then dates.
The "|" might be table borders.
But the data lines after that are each cell content.
Let's assume the table has 9 columns? Header cells: 1. OFFICE, 2. NAME, 3. Date of Authority, 4. Annual Salary, 5. Appointment, 6. House or Quarters..., 7. Leave taken during 1932, 8. Date of First Appointment. That's 8 columns. But there is a "|" before FIRE BRIGADE and before 224, maybe indicating a row header or something.
The data lines: starting from "Engineer Class III," that's first cell of first data row (Office). Then "Do." second row Office? Or second cell of first row? Since each line is a cell, we need to know the number of columns to map lines to rows.
We have 18 rows (from 18 names). 8 columns => 144 cells. But the data lines count: let's count lines from "Engineer Class III," to end.
I'll count roughly:
1 Engineer Class III,
2 Do.,
3 Wong Tsung.
4 Yeung Ling,
5 1st September,
6 1925.
7 Establishment.
8 $516
9 Quarters.
10 1st June,
11 Do.
12 469
13 Do.
14 1931.
15 Do.,
16 Chau Chi-ling.
17 1st January,
18 Do.
19 516
20 Do.
21 1927.
22 Do.,
23 Leung Ki,
24 1st May,
25 Do.
26 468
27 Do.
28 1931.
29 Do.,
30 Tsang Yin.
31 1st January,
32 Do.
33 468
34 Do.
35 1932.
36 Do.,
37 Li Yan-kwo.
38 Do.
39 Do.
40 168
41 Do.
42 Engine Driver,
43 Fan Lam.
44 1st April,
45 Do.
46 384
47 Do.
48 1928.
49 Do.,
50 Lau Chenng.
51 Do.
52 Do.
53 372
54 Do.
55 Do.,
56 Chong Kan.
57 1st July,
58 Do.
59 384
60 Do.
61 1922.
62 Do.,
63 Ip Fat.
64 1st March,
65 Do.
66 348
67 Do.
68 1930.
69 Do.
70 Do
71 Yeung Yeung.
72 1st May,
73 Do.
74 324
75 Do.
76 1932.
77 Wong Fn.
78 15th October,
79 Do.
80 381
81 Do.
82 1932.
83 Do.,
84 Leung Sai-lo.
85 17th June,
86 Do.
87 504
88 Do.
89 1923.
90 Do.,
91 Leung Kam-sni.
92 1st January,
93 Do.
94 396
95 Do.
96 1927.
97 Do.,
98 Do.,
99 Do.,
100 Au Củo-ling.
101 Tang Kau.
102 1st January,
103 Do.
104 348
105 Do.
106 ↑
107 1930.
108 1st June,
109 Do.
110 336
111 Do.
112 1931.
113 Kwok Nguu.
114 15th July,
115 Do.
116 396
117 Do.
118 1981.
119 Do..
120 Wong Taeung-keung.
121 23rd November,
122 Do.
123 396
124 Do.
125 ARTI
126 1982.
That's 126 lines. 126 / 18 = 7. So maybe 7 columns? But header suggests 8. Let's see: 126 cells for 18 rows => 7 columns exactly. 18*7=126. Good.
So the table has 7 columns, not 8. The header might have 7 columns. Let's check header lines:
1 OFFICE.
2 NAME.
3 of
4 Authority.
5 Aummal Salary.
6 Appointment,
7 House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
8 Leave taken
9 Juring 1932.
10 |
11 FIRE BRIGADE,—Continued.
12 Date of First Appointment.
13 |
14 224
15 (J 124 )
That's 15 lines before dates. But the dates list might be part of the table? Actually the dates list might be the "Date of First Appointment" column for each row, but they appear before the data rows. In the line-by-line cell output, the dates list might be the cells of the last column for all rows, but placed before the other columns? That would be column-major order for that column.
But the data lines we counted (126) start after the dates list. The dates list has 18 dates (each date might be multiple lines). Let's count date lines: from "11th January, 1920," to "23rd November, 1982." That's many lines. But they are not in the 126 count.
If the table is 18 rows x 7 columns = 126 cells, and we have 126 data lines after the dates, then the dates list might be separate (maybe the Date of First Appointment column was read first).
But the header includes "Date of First Appointment." as a column. In the line-by-line output, the header cells are also listed line by line. The header lines before the first "|" might be the first row of the table (header row). But they are split across lines.
Let's parse the header as a row of cells. The lines before the first "|" (which appears after "Juring 1932.") are:
1 OFFICE.
2 NAME.
3 of
4 Authority.
5 Aummal Salary.
6 Appointment,
7 House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
8 Leave taken
9 Juring 1932.
That's 9 lines. But we expect 7 columns. Maybe "of" and "Authority." are two cells? But "Date of Authority" is one column. Could be "Date" missing? The OCR might have missed "Date". The line "of" might be "Date of" but "Date" not captured. "Authority." is separate.
Similarly, "Leave taken" and "Juring 1932." might be two cells? But "Leave taken during 1932" is one column.
The "|" might indicate a new row or a separator.
After the first "|", we have "FIRE BRIGADE,—Continued." and "Date of First Appointment." and another "|". That might be a second header row? Or a title row.
Then "224" and "(J 124 )" might be page number.
Then the dates list: 18 dates, each date might be a cell for the "Date of First Appointment" column for each row. But they appear before the data rows.
Then the data rows follow, each row having 7 cells (since 126 lines / 18 = 7). But the data lines we listed start with "Engineer Class III," which would be the first cell of first data row (Office). Then "Do." second cell? But second cell should be Name. But "Do." is not a name. So maybe the data rows have 7 cells but the columns are different.
Let's map the 7 columns based on the data.
We have 18 rows. For each row, we have 7 cells in order.
Row1 cells (lines 1-7):
Row2 cells (lines 8-14):
Row3 cells (lines 15-21):
Row4 cells (lines 22-28):
Row5 cells (lines 29-35):
Row6 cells (lines 36-42):
Row7 cells (lines 43-49):
Row8 cells (lines 50-56):
Row9 cells (lines 57-63):
Row10 cells (lines 64-70):
Row11 cells (lines 71-77):
Row12 cells (lines 78-84):
Row13 cells (lines 85-91):
Row14 cells (lines 92-98):
Row15 cells (lines 99-105):
Row16 cells (lines 106-112):
Row17 cells (lines 113-119):
Row18 cells (lines 120-126):
This mapping gives 7 cells per row. But the content seems misaligned: Row1 has "Engineer Class III", "Do.", "Wong Tsung", "Yeung Ling", "1st September", "1925", "Establishment". That's 7 cells. But "Do." in second cell? That would be the second column. If columns are: 1. Office, 2. Name, 3. Date of Authority, 4. Annual Salary, 5. Appointment, 6. Quarters, 7. Leave? But then row1 has two names? "Wong Tsung" and "Yeung Ling". And "Do." in column2 is weird.
Maybe the table has a multi-row header or the first row is not a data row but a sub-header.
Notice "Engineer Class III" might be a section header, not a data row. In the original table, there might be a row for "Engineer Class III" as a grade, then the names under it.
But the OCR reads line by line, so "Engineer Class III" is a cell in the first column of the first data row. Then "Do." in the second column? That doesn't make sense.
Perhaps the table has 8 columns, and the first column is "Office", second "Name", third "Date of Authority", fourth "Annual Salary", fifth "Appointment", sixth "House/Quarters", seventh "Leave", eighth "Date of First Appointment". But we have 7 cells per row in the data lines. However, the dates list (18 dates) might be the eighth column, read separately.
If we combine the dates list as the eighth column for each row, then we have 8 columns.
The dates list in order:
Now, the data rows (7 cells each) plus this date as 8th column.
But the data rows as segmented above don't align well with names.
Let's look at the names we expect: 18 names. In the segmented rows, names appear in various cells.
Row1 cell3: Wong Tsung
Row1 cell4: Yeung Ling
Row2: no name
Row3 cell2: Chau Chi-ling
Row4 cell2: Leung Ki
Row5 cell2: Tsang Yin
Row6 cell2: Li Yan-kwo
Row7 cell1: Fan Lam
Row8 cell1: Lau Chenng
Row9 cell6: Ip Fat? Actually row9 cell6 is "Do.," cell7 is "Ip Fat." Wait row9 cells: 57. 1st July, 58. Do. 59. 384 60. Do. 61. 1922. 62. Do., 63. Ip Fat. So cell7 is Ip Fat.
Row10 cell6: "Do" (line70) and cell7? Row10 has 7 cells: lines64-70: 64. 1st March, 65. Do. 66. 348 67. Do. 68. 1930. 69. Do. 70. Do. So no name in row10.
Row11 cell1: Yeung Yeung.
Row11 cell7: Wong Fn.
Row12 cell7: Leung Sai-lo.
Row13 cell7: Leung Kam-sni.
Row14 cell7: Do. (line98)
Row15 cell2: Au Củo-ling. cell3: Tang Kau.
Row16: no clear name.
Row17 cell1: Kwok Nguu.
Row18 cell1: Wong Taeung-keung.
This is messy. The segmentation into 7-cell rows might be wrong because the OCR lines might not correspond exactly to cells; some cells might span multiple lines (like "House or Quarters..." is one cell but multiple lines in header). In data, a cell might be multiple lines? But the data lines seem each line is a cell.
Maybe the table has 9 columns? 126/18=7 exactly, so 7 columns.
Let's examine the header again. The header lines before the first "|" are 9 lines. But two of them might be continuations: "of" and "Authority." belong together as "Date of Authority". "Leave taken" and "Juring 1932." belong together as "Leave taken during 1932". So that reduces to 7 header cells:
That's 7 columns. The "Date of First Appointment" appears after the "|" as a separate header? But there is a "|" then "FIRE BRIGADE,—Continued." then "Date of First Appointment." then "|". That might be a title row and a column header for the last column? But we have only 7 columns in data.
Maybe the "Date of First Appointment" is not a column in this table? But the header includes it. The OCR shows "Date of First Appointment." after the second "|". The first "|" might be a vertical line in the table separating the main columns from a remarks column? Or the table has two sections.
Look at the original OCR: "Leave taken Juring 1932. | FIRE BRIGADE,—Continued. Date of First Appointment. | 224 (J 124 )". The "|" might be table borders. So the table has columns up to "Leave taken during 1932", then a border, then "FIRE BRIGADE,—Continued." might be a row header? And "Date of First Appointment." might be another column? But then "224 (J 124 )" might be page number.
Alternatively, the table might be continued from previous page, and "FIRE BRIGADE,—Continued." is a center heading. The "Date of First Appointment" might be the last column of the table.
Given the data lines have 7 cells per row, and the header has 7 meaningful columns (if we combine split lines), I'll assume the table has 7 columns: Office, Name, Date of Authority, Annual Salary, Appointment, Quarters/Allowances, Leave taken during 1932. The Date of First Appointment might be in a separate column but not captured in the data lines? But the dates list is there.
Wait, the dates list has 18 entries. If the table has 18 rows, and Date of First Appointment is a column, then the data lines should have 8 cells per row. But we have 126 lines, 126/18=7. So either the Date of First Appointment is not in the data lines (it's in the dates list), or the data lines include it but we miscounted.
Let's count the data lines again carefully. The data lines start after "23rd November, 1982." The next line is "Engineer Class III,". Let's number from there to end.
I'll copy the data lines as a list:
Yes, 126 lines.
Now, if the table has 7 columns, then each row has 7 cells. 18 rows * 7 = 126. So the mapping of lines to rows is sequential: lines 1-7 row1, 8-14 row2, etc.
But the content of row1: cells:
Row2:
Row3:
Row4:
Row5:
Row6:
Row7:
Row8:
Row9:
Row10:
Row11:
Row12:
Row13:
Row14:
Row15:
Row16:
Row17:
Row18:
Now, we need to assign column headers to these 7 columns.
From the header lines (combined):
Col1: OFFICE
Col2: NAME
Col3: Date of Authority
Col4: Annual Salary
Col5: Appointment
Col6: House or Quarters, and Allowances...
Col7: Leave taken during 1932
But the data in row1 col1 is "Engineer Class III" (an office), col2 is "Do." (ditto for office? but col2 is Name). That doesn't match.
Maybe the columns are different. Let's look at row1: col1: Engineer Class III (office), col2: Do. (maybe ditto for office? but then col3: Wong Tsung (name), col4: Yeung Ling (name?), col5: 1st September (date), col6: 1925 (year), col7: Establishment (appointment). That suggests the table might have two name columns? Or the first row is a header for a group.
In many civil service lists, there is a row for the grade "Engineer Class III" and then the officers listed below with their names. But the table structure might have the grade in the first column, and the name in the second column, but the grade row might have the grade in first column and the rest blank or "Do.".
But here row1 has "Engineer Class III" in col1, "Do." in col2, "Wong Tsung" in col3, "Yeung Ling" in col4. That's odd.
Perhaps the OCR has misordered the columns. The instruction says: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense."
The OCR read the table line by line, but the table might have been arranged with columns not strictly left-to-right in the OCR output. However, the OCR output lines seem to be in reading order (left to right, top to bottom) for each cell. But the table might have a complex layout with merged cells.
Given the difficulty, perhaps the best is to reconstruct the table as it appears in the original by using the known structure of Hong Kong Civil Service Lists.
I recall that the Hong Kong Blue Book / Civil Service List has tables with columns: Office, Name, Date of Appointment, Annual Salary, Allowances, etc.
But here we have a specific snippet.
Given the user expects a proofread text in Markdown, and the input is OCR output, I should output a cleaned-up version of the text, preserving the table structure.
Since the OCR output is essentially a list of cells in order, I can try to reconstruct the table by assuming the first few lines are header, then the dates list is the "Date of First Appointment" column, then the data lines are the other columns.
But the data lines have 7 cells per row, and we have 18 rows. The dates list has 18 dates. So if we take the dates list as column 8, we have 8 columns.
But the header has 7 columns before the "|". The "Date of First Appointment" appears after the "|". So maybe the table has 8 columns: the first 7 as per header, and the 8th is "Date of First Appointment". The OCR read the 8th column first (the dates list), then the other 7 columns for each row.
That would make sense: The OCR read the table column by column? But the data lines are 126, which is 18*7. So the data lines are the first 7 columns for all rows, in row-major order. The dates list is the 8th column, read separately (maybe because it was at the rightmost and OCR read it first).
Thus, the original table has 8 columns. The header row has 8 cells: OFFICE, NAME, Date of Authority, Annual Salary, Appointment, House/Quarters/Allowances, Leave taken during 1932, Date of First Appointment.
The OCR output: first, the header cells for the first 7 columns (split across lines), then a "|" (border), then "FIRE BRIGADE,—Continued." (maybe a title), then "Date of First Appointment." (the 8th header), then "|", then page number, then the 8th column data (18 dates), then the first 7 columns data for all rows (126 cells).
That matches!
So we have:
But the 126 lines are in row-major order for the first 7 columns. However, the first 7 columns of row1 are lines 1-7, row2 lines 8-14, etc.
But row1 col1 is "Engineer Class III", col2 "Do.", col3 "Wong Tsung", col4 "Yeung Ling", col5 "1st September", col6 "1925", col7 "Establishment". That would be the first 7 columns for the first row. But the first row should correspond to the first date (11th Jan 1920). The name should be Wong Tsung? But col3 is Wong Tsung, col4 is Yeung Ling. That suggests that the first row actually contains two names? Or the table has two rows for the first two persons in the first row? That doesn't make sense.
Maybe the data lines are not row-major for the first 7 columns, but column-major? Let's check: If the 126 lines are column-major for the first 7 columns, then the first 18 lines would be column1 for all rows, next 18 lines column2, etc. 7 columns * 18 rows = 126. So lines 1-18: column1 (Office), lines 19-36: column2 (Name), lines 37-54: column3 (Date of Authority), lines 55-72: column4 (Annual Salary), lines 73-90: column5 (Appointment), lines 91-108: column6 (Quarters), lines 109-126: column7 (Leave).
Let's test this hypothesis.
Lines 1-18 (Office column):
That doesn't look like all offices. Line3 is "Wong Tsung." which is a name, not office. So not column-major.
Maybe the data lines are in the order of the original table reading order (left to right, top to bottom) but the table has a different number of columns per row? Or the table has merged cells.
Given the time, I think the most practical is to output the text as a cleaned-up version of the OCR, preserving the table structure as best as possible, using Markdown table with the columns as per header, and filling in the data from the OCR text in a logical way.
But the user said: "Your most important task is to fix the text flow while preserving the original paragraph structure." This is a table, not paragraphs. So I should reconstruct the table.
I will create a Markdown table with the 8 columns. I'll use the dates list for the last column. For the first 7 columns, I'll try to parse the 126 lines into 18 rows of 7 columns by assuming the OCR read the table row by row, but the first row might be a header row for the group "Engineer Class III". Actually, the table might have a row for "Engineer Class III" spanning multiple columns? But the OCR reads cells linearly.
Let's look at the original OCR text again: It starts with "OFFICE. NAME. of Authority. Aummal Salary. Appointment, House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose. Leave taken Juring 1932. | FIRE BRIGADE,—Continued. Date of First Appointment. | 224 (J 124 )". This suggests the table has a header row with those columns, then a title "FIRE BRIGADE,—Continued." maybe centered, then "Date of First Appointment" might be a column header for the last column? But it's after a "|".
Maybe the table is split into two parts: left part with columns OFFICE to Leave taken, right part with Date of First Appointment. The "|" indicates a vertical line.
In the data, the dates list appears after the page number, then the left part data.
The left part data (126 lines) might be the left part columns for all rows. The left part has 7 columns (OFFICE to Leave taken). 18 rows * 7 = 126. Good.
The right part (Date of First Appointment) has 18 dates.
So the full table has 8 columns. The left 7 columns data is in the 126 lines in row-major order. The rightmost column data is in the dates list in order.
Now, we need to split the 126 lines into 18 rows of 7 columns each. We already did that. But the content of row1 (lines1-7) seems to contain data for two persons? Let's see row1:
Col1: Engineer Class III
Col2: Do.
Col3: Wong Tsung
Col4: Yeung Ling
Col5: 1st September
Col6: 1925
Col7: Establishment
But the left part columns are: 1. OFFICE, 2. NAME, 3. Date of Authority, 4. Annual Salary, 5. Appointment, 6. House/Quarters/Allowances, 7. Leave taken during 1932.
So for row1, OFFICE = "Engineer Class III", NAME = "Do."? That can't be. NAME should be a name. "Do." means ditto, so same as previous row's name. But there is no previous row. So maybe the first row is not a data row but a sub-heading. In many tables, the grade "Engineer Class III" is written in the OFFICE column for the first officer, and then "Do." for subsequent officers. But here the first officer is Wong Tsung. So row1 should be: OFFICE: Engineer Class III, NAME: Wong Tsung, Date of Authority: ?, Annual Salary: ?, Appointment: ?, Quarters: ?, Leave: ?.
But in our segmentation, row1 has NAME = "Do." and then Wong Tsung appears in col3. So the columns are shifted.
Perhaps the table has an extra column? Or the "Do." in col2 is actually the Date of Authority for the previous? No.
Let's look at the raw lines again. The first few lines after the dates list:
"Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment."
If we consider that the table might have been read with a different column order: maybe the OCR read the table in a different order because of the
!
Dato
OFFICE.
NAME.
of
Authority.
Aummal Salary.
Appointment,
House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
Leave taken
Juring 1932.
FIRE BRIGADE,—Continued.
Date of First Appointment.
224
(J 124 )
11th January, 1920, 21st October,
31st January,
1922.
1st December, 1921. 1st May, 1927.
1st January, 1925.
17th November.
1st April, 1926.
5th November, !928.
15th Septemer, 1927. 15th July,
1931.
23rd November,
1982.
Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment.
$516
Quarters.
1st June,
Do.
469
Do.
1931.
Do.,
Chau Chi-ling.
1st January,
Do.
516
Do.
1927.
Do.,
Leung Ki,
1st May,
Do.
468
Do.
1931.
Do.,
Tsang Yin.
1st January,
Do.
468
Do.
1932.
Do.,
Li Yan-kwo.
Do.
Do.
168
Do.
Engine Driver,
Fan Lam.
1st April,
Do.
384
Do.
1928.
Do.,
Lau Chenng.
Do.
Do.
372
Do.
Do.,
Chong Kan.
1st July,
Do.
384
Do.
1922.
Do.,
Ip Fat.
1st March,
Do.
348
Do.
1930.
Do.
Do
Yeung Yeung.
1st May,
Do.
324
Do.
1932.
Wong Fn.
15th October,
Do.
381
Do.
1932.
Do.,
Leung Sai-lo.
17th June,
Do.
504
Do.
1923.
Do.,
Leung Kam-sni.
1st January,
Do.
396
Do.
1927.
Do.,
Do.,
Do.,
Au Củo-ling.
Tang Kau.
1st January,
Do.
348
Do.
↑
1930.
1st June,
Do.
336
Do.
1931.
Kwok Nguu.
15th July,
Do.
396
Do.
1981.
Do..
Wong Taeung-keung.
23rd November,
Do.
396
Do.
ARTI
1982.
!
Dato
OFFICE.
NAME.
of
Authority.
Aummal Salary.
Appointment,
House or Quarters, and Allowancea for Rent, Entertainment, Personal, or for any other purpose.
Leave taken
Juring 1932.
FIRE BRIGADE,—Continued.
Date of First Appointment.
224
(J 124 )
11th January, 1920, 21st October,
31st January,
1922.
1st December, 1921. 1st May, 1927.
1st January, 1925.
17th November.
1st April, 1926.
5th November, !928.
15th Septemer, 1927. 15th July,
1931.
23rd November,
1982.
Engineer Class III,
Do.,
Wong Tsung.
Yeung Ling,
1st September,
1925.
Establishment.
$516
Quarters.
1st June,
Do.
469
Do.
1931.
Do.,
Chau Chi-ling.
1st January,
Do.
516
Do.
1927.
Do.,
Leung Ki,
1st May,
Do.
468
Do.
1931.
Do.,
Tsang Yin.
1st January,
Do.
468
Do.
1932.
Do.,
Li Yan-kwo.
Do.
Do.
168
Do.
Engine Driver,
Fan Lam.
1st April,
Do.
384
Do.
1928.
Do.,
Lau Chenng.
Do.
Do.
372
Do.
Do.,
Chong Kan.
1st July,
Do.
384
Do.
1922.
Do.,
Ip Fat.
1st March,
Do.
348
Do.
1930.
Do.
Do
Yeung Yeung.
1st May,
Do.
324
Do.
1932.
Wong Fn.
15th October,
Do.
381
Do.
1932.
Do.,
Leung Sai-lo.
17th June,
Do.
504
Do.
1923.
Do.,
Leung Kam-sni.
1st January,
Do.
396
Do.
1927.
Do.,
Do.,
Do.,
Au Củo-ling.
Tang Kau.
1st January,
Do.
348
Do.
↑
1930.
1st June,
Do.
336
Do.
1931.
Kwok Nguu.
15th July,
Do.
396
Do.
1981.
Do..
Wong Taeung-keung.
23rd November,
Do.
396
Do.
ARTI
1982.
No comments yet.
Private notes are available after approval.