The user wants me to proofread OCR output of a historical document about Hong Kong police pensions. The text appears to be a table with columns: Name of Pensioner, Amount, Authority, Date from which the Pension has been paid, Amount of Emolument when last employed in Public Service, Present Age of Pensioner, Cause of Retirement.
The OCR is messy with broken lines, misaligned columns, garbled characters. I need to reconstruct the table properly in Markdown.
First, I need to parse the data. The original seems to have columns: Name, Amount (£ s. d.), Authority, Date, Emolument, Age, Cause.
But the OCR has mixed up lines. Let me try to reconstruct row by row.
Looking at the text:
First line: "324" maybe page number.
"(L 20)" maybe reference.
"POLICE PENSIONS." header.
Then column headers: "Name of Pensioner.", "Amount.", "Amount. Authority.", "Date from which the Pension Ins been paid.", "Amount of Emolument when last employed in Public Service.", "Present Age of Pensioner.", "Cause of Retirement,"
Then "£ S. d." maybe subheader for Amount.
Then rows:
"Delail Khan, 67.21 16th Nov., 1910. 210,00 51 Medical Cer- tificate."
But Amount column seems to have £ s. d. format. "67.21" maybe £67 2s 1d? But there is "16th Nov., 1910." under Date? Actually "16th Nov., 1910." appears after 67.21. Then "210,00" maybe emolument? Then "51" age, "Medical Certificate" cause.
But the columns are misaligned. Let's examine the raw text lines:
Delail Khan,
67.21
16th Nov., 1910.
210,00
51
Medical Cer-
tificate.
Albert Wilson,
288.00
1st Jan., 1911.
$64,00
William Robertson,
831,00
10th June, 1911.
1,662.00
Amar Singh,
115,00
11th Aug., 1911.
300,00
Su Toi,
180.00
14th Aug., 1911.
270,00
Hussein Shah,.
114.00
Chu Ping,
81,00
Mahomed Shah,
64.82
Makbool Shah,
65.65
Anokh Singh,.
52.70
Ordinance No. 11 of 1900.
19th Nov.. 1911.
342.00
* 8 € 2 8
11
65
63
On Expiration of
Term of Service.
75
争
60
"
1st Jan,
1912.
180.00
53
11
1st Mar., 1912,
210.00
62
Do.
210.00
63
16th Mar., 1912.
186.00
8
Injured on duty.
On Expiration of
Term of service.
50
די
Kesur Singh,
* 4 4 4 4
62.95
1st April, 1912.
186.00
64
1
Kadar Bux,.......
52.70
Do.
186.00
46
Lall Singh,
Munga Khan,
Bikh Mahomed,
62.00
1st June, 1912.
186.00
61
Medical Cer- tificate.
¡On Expiration of Term of Service.
Medical Cer- tificate.
60.06
16th July, 1912.
216.00
54
On Expiration of
Term of Service.
$2.70
16th Aug,, 1912.
186.00
49
Medical Cer- tificate.
Kam Kour (widow of
Singh) Rs. 36,
P.C. 769 Jhanda
2 8 0
Granted Special Pension by The Governor in Connell on the death of lier husband who was killed on duty,
19th Ang., 1912,
39
Lau Fuk.....
60.00
1st Sept., 1912.
180,00
50
| Medical Cer-
tificate.
Chan Fan,
Ahmed Deen,
110.19
14th Jan, 1913.
300.00
56
55.80
1st April, 1912.
186,00
པ་
On Expiration of „Term of Service.
Medical Cer- tificate.
Atta Mohamed,
47.12
Do.
186,00
47
**
Nathy Khan, ......
49.60
Do.
186,00
42
19
Henry George Baker,
1,178.57
Li Pui,
60.00
Abdullah,
Mehr Khan,
74.64
55.80
John Lenaghan,
42 0 0
Mahomed Shab,
5.270
Ordinance No. 11 of 1900.
10th April, 1913. 2,016.00
65
On Expiration of
1st July, 1913.
180.00
53
Term of Service.
Medical Certificate.
16th July, 1913.
216.00
48
12
1st Oct., 1913,
186,00
*
48
11
24th Oct,, 1913, £145,0.0 52
29
Gunda Singh,
74.78
Jhande Khun,
136.50
A
William Camerou,
746.76
1st Jan, 1914.
1st Mar., 1914.
1st May, 1914.
23rd May, 1914.
$186.00
55
"
246.00
62
On Expiration of
Terms of Service.
360.00
48
Medical Certificate.
1,632.00
64
fn Expiration of
Term of Service.
Karm Elahi,
95.86
Alfred George Dymond
676.80
1st June, 1914.
6th June, 1914.
246.00
1,632.00
62
29
59
>1
This is a mess. It seems the OCR read columns vertically? Actually the original document likely had a table with multiple columns. The OCR may have read line by line across columns, causing interleaving.
We need to reconstruct the table logically. Let's think about the typical columns for a pension list:
The OCR output seems to have lost column boundaries. We need to parse each row.
Let's try to identify rows by looking for names and dates.
First row: "Delail Khan, 67.21 16th Nov., 1910. 210,00 51 Medical Cer- tificate."
That seems plausible: Name: Delail Khan, Amount: £67 2s 1d? Actually "67.21" could be £67.21? But old currency: pounds, shillings, pence. Usually written as £67 2s 1d. The OCR might have merged. "67.21" might be 67.21? But there is "£ S. d." header. So amount column likely has three subcolumns. The OCR may have combined them.
Second: "Albert Wilson, 288.00 1st Jan., 1911. $64,00" Wait, $64,00? That might be emolument in dollars? But the column "Amount of Emolument when last employed in Public Service" might be in dollars. The pension amount might be in pounds. So "288.00" could be pension amount £288? Then date 1st Jan 1911, emolument $64.00? But then missing age and cause.
Third: "William Robertson, 831,00 10th June, 1911. 1,662.00" Pension £831? Emolument 1,662? Age? Cause?
Fourth: "Amar Singh, 115,00 11th Aug., 1911. 300,00"
Fifth: "Su Toi, 180.00 14th Aug., 1911. 270,00"
Sixth: "Hussein Shah,. 114.00" then next line "Chu Ping, 81,00" then "Mahomed Shah, 64.82" then "Makbool Shah, 65.65" then "Anokh Singh,. 52.70" then "Ordinance No. 11 of 1900. 19th Nov.. 1911. 342.00" This seems like a new row: Authority "Ordinance No. 11 of 1900", Date "19th Nov. 1911", Emolument "342.00"? But missing name? Actually maybe the previous names (Hussein Shah, Chu Ping, Mahomed Shah, Makbool Shah, Anokh Singh) are all under the same authority and date? Or they are separate rows but the OCR merged lines.
Look at the pattern: After "Anokh Singh,. 52.70" there is "Ordinance No. 11 of 1900. 19th Nov.. 1911. 342.00". Then " 8 € 2 8 11 65 63 On Expiration of Term of Service. 75 争 60 " 1st Jan, 1912. 180.00 53 11 1st Mar., 1912, 210.00 62 Do. 210.00 63 16th Mar., 1912. 186.00 8 Injured on duty. On Expiration of Term of service. 50 די Kesur Singh, 4 4 4 4 62.95 1st April, 1912. 186.00 64 1 Kadar Bux,....... 52.70 Do. 186.00 46 Lall Singh, Munga Khan, Bikh Mahomed, 62.00 1st June, 1912. 186.00 61 Medical Cer- tificate. ¡On Expiration of Term of Service. Medical Cer- tificate. 60.06 16th July, 1912. 216.00 54 On Expiration of Term of Service. $2.70 16th Aug,, 1912. 186.00 49 Medical Cer- tificate. Kam Kour (widow of Singh) Rs. 36, P.C. 769 Jhanda 2 8 0 Granted Special Pension by The Governor in Connell on the death of lier husband who was killed on duty, 19th Ang., 1912, 39 Lau Fuk..... 60.00 1st Sept., 1912. 180,00 50 | Medical Cer- tificate. Chan Fan, Ahmed Deen, 110.19 14th Jan, 1913. 300.00 56 55.80 1st April, 1912. 186,00 פָּ On Expiration of „Term of Service. Medical Cer- tificate. Atta Mohamed, 47.12 Do. 186,00 47 * Nathy Khan, ...... 49.60 Do. 186,00 42 19 Henry George Baker, 1,178.57 Li Pui, 60.00 Abdullah, Mehr Khan, 74.64 55.80 John Lenaghan, 42 0 0 Mahomed Shab, 5.270 Ordinance No. 11 of 1900. 10th April, 1913. 2,016.00 65 On Expiration of 1st July, 1913. 180.00 53 Term of Service. Medical Certificate. 16th July, 1913. 216.00 48 12 1st Oct., 1913, 186,00 48 11 24th Oct,, 1913, £145,0.0 52 29 Gunda Singh, 74.78 Jhande Khun, 136.50 A William Camerou, 746.76 1st Jan, 1914. 1st Mar., 1914. 1st May, 1914. 23rd May, 1914. $186.00 55 " 246.00 62 On Expiration of Terms of Service. 360.00 48 Medical Certificate. 1,632.00 64 fn Expiration of Term of Service. Karm Elahi, 95.86 Alfred George Dymond 676.80 1st June, 1914. 6th June, 1914. 246.00 1,632.00 62 29 59 >1"
This is extremely messy. It appears the OCR read the table column by column? Or the original had multiple columns per page and OCR read across columns.
Given the complexity, perhaps the best approach is to reconstruct the table as best as possible, using logical grouping. However, the user expects a proofread text in Markdown, preserving paragraph breaks and tables. Since the original is a table, we should output a Markdown table.
But we need to correctly parse rows. This is a challenging OCR correction task. I need to use my knowledge of typical Hong Kong police pension records. The document appears to be a list of police pensioners with details.
Let's try to parse systematically. The OCR text seems to have lines that correspond to rows but with missing fields. Perhaps the original table had columns: Name, Pension Amount (£ s. d.), Authority, Date of Pension, Last Emolument, Age, Cause.
The OCR output might have lost the column separation, causing fields to run together. But we can try to identify each row by the presence of a name and a date.
Let's list all names that appear:
Also there are some entries like "Ordinance No. 11 of 1900" which is an authority, not a name.
Now, each row should have a name, pension amount, authority, date, emolument, age, cause.
Let's try to match each name with surrounding data.
Start from top:
"Delail Khan, 67.21 16th Nov., 1910. 210,00 51 Medical Cer- tificate."
Interpretation:
But there is an "Authority" column. Could be "Ordinance No. 11 of 1900" for many. But not shown here.
Next: "Albert Wilson, 288.00 1st Jan., 1911. $64,00"
Next: "William Robertson, 831,00 10th June, 1911. 1,662.00"
Next: "Amar Singh, 115,00 11th Aug., 1911. 300,00"
Next: "Su Toi, 180.00 14th Aug., 1911. 270,00"
Next: "Hussein Shah,. 114.00" then new line "Chu Ping, 81,00" then "Mahomed Shah, 64.82" then "Makbool Shah, 65.65" then "Anokh Singh,. 52.70" then "Ordinance No. 11 of 1900. 19th Nov.. 1911. 342.00"
This suggests that Hussein Shah, Chu Ping, Mahomed Shah, Makbool Shah, Anokh Singh are all pensioners with amounts 114.00, 81.00, 64.82, 65.65, 52.70 respectively, and they share the same authority "Ordinance No. 11 of 1900", date "19th Nov. 1911", emolument "342.00"? But emolument 342.00 for all? That seems odd. Maybe the emolument is 342.00 for the first? Or maybe the line "Ordinance No. 11 of 1900. 19th Nov.. 1911. 342.00" belongs to Anokh Singh? But then what about the others? They might have different dates and emoluments.
Look at the next lines: " 8 € 2 8 11 65 63 On Expiration of Term of Service. 75 争 60 " 1st Jan, 1912. 180.00 53 11 1st Mar., 1912, 210.00 62 Do. 210.00 63 16th Mar., 1912. 186.00 8 Injured on duty. On Expiration of Term of service. 50 די Kesur Singh, 4 4 4 4 62.95 1st April, 1912. 186.00 64 1 Kadar Bux,....... 52.70 Do. 186.00 46 Lall Singh, Munga Khan, Bikh Mahomed, 62.00 1st June, 1912. 186.00 61 Medical Cer- tificate. ¡On Expiration of Term of Service. Medical Cer- tificate. 60.06 16th July, 1912. 216.00 54 On Expiration of Term of Service. $2.70 16th Aug,, 1912. 186.00 49 Medical Cer- tificate. Kam Kour (widow of Singh) Rs. 36, P.C. 769 Jhanda 2 8 0 Granted Special Pension by The Governor in Connell on the death of lier husband who was killed on duty, 19th Ang., 1912, 39 Lau Fuk..... 60.00 1st Sept., 1912. 180,00 50 | Medical Cer- tificate. Chan Fan, Ahmed Deen, 110.19 14th Jan, 1913. 300.00 56 55.80 1st April, 1912. 186,00 פָּ On Expiration of „Term of Service. Medical Cer- tificate. Atta Mohamed, 47.12 Do. 186,00 47 * Nathy Khan, ...... 49.60 Do. 186,00 42 19 Henry George Baker, 1,178.57 Li Pui, 60.00 Abdullah, Mehr Khan, 74.64 55.80 John Lenaghan, 42 0 0 Mahomed Shab, 5.270 Ordinance No. 11 of 1900. 10th April, 1913. 2,016.00 65 On Expiration of 1st July, 1913. 180.00 53 Term of Service. Medical Certificate. 16th July, 1913. 216.00 48 12 1st Oct., 1913, 186,00 48 11 24th Oct,, 1913, £145,0.0 52 29 Gunda Singh, 74.78 Jhande Khun, 136.50 A William Camerou, 746.76 1st Jan, 1914. 1st Mar., 1914. 1st May, 1914. 23rd May, 1914. $186.00 55 " 246.00 62 On Expiration of Terms of Service. 360.00 48 Medical Certificate. 1,632.00 64 fn Expiration of Term of Service. Karm Elahi, 95.86 Alfred George Dymond 676.80 1st June, 1914. 6th June, 1914. 246.00 1,632.00 62 29 59 >1"
This is extremely garbled. It seems the OCR has mixed multiple columns. Perhaps the original table had two columns per page? Or the OCR read the table in a zigzag.
Given the difficulty, maybe the best is to output the text as a cleaned-up version preserving the table structure as much as possible, but given the instruction to reconstruct sentences and paragraphs, and the fact that this is a table, we should output a Markdown table with the columns we can deduce.
But we need to be accurate. However, the user is asking to proofread OCR output. They provided the OCR text. They expect us to return the proofread text in Markdown. We must do our best.
Given the complexity, I might need to make assumptions and reconstruct the table row by row based on typical patterns.
Let's try to parse the data as a list of records. I'll write a script in my mind to parse.
But as an AI, I can attempt to manually reconstruct.
First, note the page number "324" and "(L 20)" maybe a reference.
The header: "POLICE PENSIONS."
Then column headers: "Name of Pensioner.", "Amount.", "Amount. Authority.", "Date from which the Pension Ins been paid.", "Amount of Emolument when last employed in Public Service.", "Present Age of Pensioner.", "Cause of Retirement,"
Then "£ S. d." likely subheader for Amount.
Now, the data rows. Let's try to split the OCR text into lines as they appear in the input. The input is a single block of text with line breaks. I'll split by newline.
Lines:
This is a line-by-line breakdown. Now, the original table likely had multiple columns. The OCR seems to have read each cell as a separate line? But there are many lines. For example, "Delail Khan," then "67.21" then "16th Nov., 1910." then "210,00" then "51" then "Medical Cer-" then "tificate." That's 7 lines for one row. That matches 7 columns. Good! So each row is spread across 7 lines (one per column). But then "Albert Wilson," "288.00" "1st Jan., 1911." "$64,00" then next line "William Robertson," ... wait, after "$64,00" we have "William Robertson," which is the next row's name. But we are missing age and cause for Albert Wilson. Actually, the columns: Name, Amount, Authority, Date, Emolument, Age, Cause. For Albert Wilson, we have Name, Amount, Date, Emolument? But missing Authority, Age, Cause. However, the line after "$64,00" is "William Robertson,". So maybe the Authority column is blank (ditto), and Age and Cause are missing? Or the OCR missed lines.
Let's check the pattern: For Delail Khan, we have 7 lines: Name, Amount, Date, Emolument, Age, Cause (split across two lines). That's 6 fields? Actually: Name, Amount, Date, Emolument, Age, Cause. That's 6 columns. But the header has 7 columns: Name, Amount, Authority, Date, Emolument, Age, Cause. So Authority is missing for Delail Khan. Maybe the Authority is "Ordinance No. 11 of 1900" but appears later.
Look at line 59: "Ordinance No. 11 of 1900." line 60: "19th Nov.. 1911." line 61: "342.00". That could be Authority, Date, Emolument for a row. But which name? Possibly for the group of Hussein Shah, Chu Ping, etc. But they each have their own amounts (lines 50,52,54,56,58). They might share the same Authority, Date, Emolument? But then they would need individual ages and causes.
The lines after 61: "* 8 € 2 8" "11" "65" "63" "On Expiration of" "Term of Service." "75" "争" "60" ... This is garbled.
Maybe the table is actually two separate tables? Or the OCR has merged two pages.
Given the difficulty, perhaps the best is to output the text as a cleaned-up version of the OCR, preserving the line breaks as paragraphs? But the instruction says to reconstruct sentences and paragraphs, and for tables, reconstruct using Markdown table syntax.
Given the data is tabular, we must produce a Markdown table. But we need to infer the correct rows.
Let's try to group lines into rows of 7 columns. Since the header has 7 columns, each row should have 7 cells. The OCR output seems to have each cell on a new line, but sometimes cells are split across lines (like "Medical Cer-" "tificate."). Also, some rows may have missing cells.
We can attempt to parse by assuming that each row starts with a name (capitalized, ends with comma). Then the next lines are the subsequent columns until the next name.
Let's list all lines that look like names (ending with comma or "....."):
Also "Ordinance No. 11 of 1900." appears twice, not a name.
Now, between names, there are lines that are data. Let's try to assign each data line to the preceding name until the next name.
But the data lines are not consistently 6 per name. For Delail Khan, we have 6 data lines (67.21, 16th Nov., 1910., 210,00, 51, Medical Cer-, tificate.) That's 6 lines for 6 columns (Amount, Authority?, Date, Emolument, Age, Cause). But Authority missing? Actually the third column is "Amount. Authority." maybe a combined column? The header says "Amount. Authority." Could be two subcolumns? Or maybe it's "Amount" and "Authority" separate. The header lines: "Amount." then "Amount. Authority." That's weird.
Let's read the header lines exactly as in OCR:
"Name of Pensioner.
Amount.
Amount. Authority.
Date from which the Pension
Ins been paid.
Amount of Emolument when last employed in Public Service.
Present Age of
Pensioner.
Cause of Retirement,"
This suggests there are 7 columns: 1. Name of Pensioner, 2. Amount, 3. Authority, 4. Date, 5. Emolument, 6. Age, 7. Cause.
But the OCR has "Amount." and "Amount. Authority." as separate lines? Actually they might be two header lines for the same column? Or maybe the table has two amount columns: one for pension amount, one for something else? The header "Amount. Authority." might be a merged cell.
Given the confusion, I think the original table likely had columns: Name, Pension Amount (£ s. d.), Authority, Date of Pension, Last Emolument, Age, Cause.
The OCR has "£ S. d." as subheader for the pension amount.
Now, for each row, we need to extract 7 fields.
Let's attempt to parse sequentially, using the fact that each field is on a new line, but some fields span multiple lines (like cause). Also, some fields may be empty.
We'll write a pseudo-parser: iterate through lines, when we see a line that looks like a name (ends with comma or "....."), start a new row. Then collect subsequent non-name lines as fields for that row until we have 6 fields (since name is first). But some rows may have fewer fields.
But the lines include many non-name lines that are not fields? Like "* 8 € 2 8", "11", "65", "63", "On Expiration of", "Term of Service.", "75", "争", "60", etc. These might be fields for the previous rows? Or they are garbled.
Maybe the OCR has read the table column by column? For example, first column (names) all listed, then second column (amounts) all listed, etc. But the lines are interleaved.
Look at the sequence: After "Anokh Singh,." and "52.70", we have "Ordinance No. 11 of 1900." which could be the Authority for the previous few names? Then "19th Nov.. 1911." Date, "342.00" Emolument. Then "* 8 € 2 8" maybe Age? "11" "65" "63" maybe ages for multiple pensioners? "On Expiration of Term of Service." Cause. "75" "争" "60" ... This is messy.
Given the time, perhaps the best is to output the text as a cleaned-up version preserving the original line breaks but fixing obvious OCR errors, and present it as a table with the columns we can identify, but note that the data is messy.
However, the user expects a proofread text. They might be testing the ability to handle OCR. I need to produce a Markdown table that represents the data as best as possible.
Given the complexity, I'll attempt to reconstruct the table by manually matching each pensioner with their data from the lines, using logical grouping.
Let's list the pensioners in order as they appear in the text (by name lines):
Now, let's try to find data for each.
We have lines with dates:
These dates likely correspond to "Date from which the Pension has been paid" for each pensioner.
Now, pension amounts (second column) appear as numbers:
67.21, 288.00, 831,00, 115,00, 180.00, 114.00, 81,00, 64.82, 65.65, 52.70, 62.95, 52.70, 62.00, 60.06, 60.00, 110.19, 55.80, 47.12, 49.60, 1,178.57, 60.00, 74.64, 55.80, 42 0 0, 5.270, 74.78, 136.50, 746.76, 95.86, 676.80.
But some of these might be emoluments.
Emoluments appear as: 210,00, $64,00, 1,662.00, 300,00, 270,00, 342.00, 180.00, 210.00, 210.00, 186.00, 186.00, 186.00, 216.00, 186.00, 180,00, 300.00, 186,00, 186,00, 186,00, 2,016.00, 180.00, 216.00, 186,00, £145,0.0, $186.00, 246.00, 360.00, 1,632.00, 246.00, 1,632.00.
Ages: 51, 65, 63, 75, 60, 53, 62, 63, 8? (maybe 58?), 50, 64, 46, 61, 54, 49, 39, 50, 56, 47, 42, 19? (maybe 65?), 53, 48, 48, 52, 29, 55, 62, 48, 64, 62, 29, 59.
Causes: Medical Certificate, On Expiration of Term of Service, Injured on duty, Granted Special Pension..., etc.
Authorities: Ordinance No. 11 of 1900, Do. (ditto), maybe "P.C. 769" for Kam Kour.
Now, let's try to match each pensioner with a date. The dates appear in chronological order. The list of pensioners seems roughly chronological by pension start date.
Let's assign dates to pensioners in order:
Let's list dates in order with possible pensioners:
This is too messy to fully parse manually in this response. However, the user likely expects a proofread version of the OCR text, not necessarily a perfect database. Since the instruction says "Return the proofread text in standard Markdown", and "Use Markdown table syntax to reconstruct tabular data", I should produce a Markdown table with the columns and rows as best as I can.
Given the difficulty, I might output a table with the columns: Name, Pension Amount, Authority, Date, Emolument, Age, Cause, and fill in the data I can clearly identify, leaving ambiguous ones with ... or best guess.
But the user said "Do not add or remove substantive words. Do not rephrase, rewrite, or change the author's style, tone, or word choice. Only correct what is clearly wrong." So I should not guess missing data; I should only correct OCR errors in the existing text. However, the OCR text is a jumble of lines. The "proofread text" should be the corrected text of the document, which is a table. So I need to reconstruct the table as it originally appeared.
Perhaps the original document is a known historical record. I could search my knowledge: "Police Pensions Hong Kong 1910-1914" maybe a Blue Book? But I don't have exact data.
Given the constraints, I'll attempt to produce a cleaned-up version of the OCR text, preserving the line structure but fixing obvious errors (like "Ins" -> "has", "Connell" -> "Council", "Ang." -> "Aug.", "פָּ" -> remove, "די" -> remove, "争" -> remove, "* 8 € 2 8" -> maybe "£8 2s 8d"?). But the instruction says to reconstruct sentences and paragraphs, and for tables, use Markdown table syntax.
I think the best is to output a Markdown table with the header and then each row as parsed from the OCR lines, using the line-by-line grouping of 7 lines per row (since header has 7 columns). Let's test that hypothesis: For Delail Khan, we have 7 lines (name, amount, date, emolument, age, cause part1, cause part2). That's 7 lines for 7 columns? Actually cause is one column but split across two lines. So maybe each column corresponds to one line, but cause sometimes wraps.
If we assume each row takes 7 lines (one per column), then we can chunk the lines after the header into groups of 7. But the header itself takes many lines. After "£ S. d." line, the data starts. Let's count lines from "Delail Khan," to end. There are 208 lines from line 26 to 233. 208 / 7 = 29.7, not integer.
Maybe the table has two columns per page? Not sure.
Given the time, I'll produce a Markdown table that includes all the legible data in a structured way, but I'll note that the OCR is fragmented. However, the user said "Output only in Markdown. Do not wrap your response in code fences or HTML tags." So I must output Markdown directly.
I'll create a table with the columns: Name, Pension Amount (£ s. d.), Authority, Date of Pension, Last Emolument, Age, Cause of Retirement.
I'll fill rows based on the most coherent sequences.
Let's attempt to parse row by row using the name lines as anchors.
I'll write a simplified reconstruction:
Row 1: Delail Khan | £67 2s 1d? | (blank) | 16th Nov. 1910 | $210.00 | 51 | Medical Certificate
Row 2: Albert Wilson | £288.00 | (blank) | 1st Jan. 1911 | $64.00 | ? | ?
Row 3: William Robertson | £831.00 | (blank) | 10th June 1911 | $1,662.00 | ? | ?
Row 4: Amar Singh | £115.00 | (blank) | 11th Aug. 1911 | $300.00 | ? | ?
Row 5: Su Toi | £180.00 | (blank) | 14th Aug. 1911 | $270.00 | ? | ?
Row 6: Hussein Shah | £114.00 | Ordinance No. 11 of 1900 | 19th Nov. 1911 | $342.00 | 65? | On Expiration of Term of Service
Row 7: Chu Ping | £81.00 | Ordinance No. 11 of 1900 | 19th Nov. 1911 | $342.00 | 63? | On Expiration of Term of Service
Row 8: Mahomed Shah | £64.82 | Ordinance No. 11 of 1900 | 19th Nov. 1911 | $342.00 | 75? | On Expiration of Term of Service
Row 9: Makbool Shah | £65.65 | Ordinance No. 11 of 1900 | 19th Nov. 1911 | $342.00 | 60? | On Expiration of Term of Service
Row 10: Anokh Singh | £52.70 | Ordinance No. 11 of 1900 | 19th Nov. 1911 | $342.00 | ? | ?
Row 11: Kesur Singh | £62.95 | Ordinance No. 11 of 1900 | 1st April 1912 | $186.00 | 64 | On Expiration of Term of Service
Row 12: Kadar Bux | £52.70 | Do. | 1st April 1912? | $186.00 | 46 | On Expiration of Term of Service
Row 13: Lall Singh | £62.00 | Do. | 1st June 1912 | $186.00 | 61 | Medical Certificate
Row 14: Munga Khan | ? | Do. | 1st June 1912 | $186.00 | ? | On Expiration of Term of Service
Row 15: Bikh Mahomed | ? | Do. | 1st June 1912 | $186.00 | ? | Medical Certificate
Row 16: Kam Kour (widow of Singh) | Rs. 36 | P.C. 769 | 19th Aug. 1912 | (special pension) | 39 | Granted Special Pension by the Governor in Council on the death of her husband who was killed on duty
Row 17: Lau Fuk | £60.00 | (blank) | 1st Sept. 1912 | $180.00 | 50 | Medical Certificate
Row 18: Chan Fan | £110.19 | (blank) | 14th Jan. 1913 | $300.00 | 56 | On Expiration of Term of Service
Row 19: Ahmed Deen | £55.80 | (blank) | 1st April 1912? | $186.00 | ? | Medical Certificate
Row 20: Atta Mohamed | £47.12 | Do. | 1st April 1912? | $186.00 | 47 | On Expiration of Term of Service
Row 21: Nathy Khan | £49.60 | Do. | 1st April 1912? | $186.00 | 42 | ?
Row 22: Henry George Baker | £1,178.57 | Ordinance No. 11 of 1900 | 10th April 1913 | $2,016.00 | 65 | On Expiration of Term of Service
Row 23: Li Pui | £60.00 | (blank) | 1st July 1913 | $180.00 | 53 | Medical Certificate
Row 24: Abdullah | £74.64 | (blank) | 16th July 1913 | $216.00 | 48 | ?
Row 25: Mehr Khan | £55.80 | (blank) | 1st Oct. 1913 | $186.00 | 48 | ?
Row 26: John Lenaghan | £42.00 | (blank) | 24th Oct. 1913 | £145.00? | 52 | ?
Row 27: Mahomed Shab | £5.270? | (blank) | ? | ? | 29 | ?
Row 28: Gunda Singh | £74.78 | (blank) | 1st Jan. 1914 | $186.00 | 55 | On Expiration of Terms of Service
Row 29: Jhande Khun | £136.50 | (blank) | 1st Mar. 1914 | $246.00 | 62 | On Expiration of Terms of Service
Row 30: William Camerou | £746.76 | (blank) | 1st May 1914 | $360.00 | 48 | Medical Certificate
Row 31: Karm Elahi | £95.86 | (blank) | 1st June 1914 | $246.00 | 62 | On Expiration of Term of Service
Row 32: Alfred George Dymond | £676.80 | (blank) | 6th June 1914 | $1,632.00 | 64? | On Expiration of Term of Service
This is speculative. But the user might not expect perfection; they want a proofread version of the OCR text. Since the OCR text is provided, maybe they want the text cleaned up but still in the same flow (i.e., not a table). However, the instruction says "Use Markdown table syntax to reconstruct tabular data." So they expect a table.
Given the ambiguity, I'll output a Markdown table with the header and as many rows as I can clearly identify, using the data from the OCR lines, correcting obvious OCR errors (like "Ins" -> "has", "Connell" -> "Council", "Ang." -> "Aug.", "Cer-" -> "Certi-", etc.). I'll also fix the hyphenation.
I'll produce a table with the columns: Name of Pensioner, Amount (£ s. d.), Authority, Date from which the Pension has been paid, Amount of Emolument when last employed in Public Service, Present Age of Pensioner, Cause of Retirement.
I'll fill rows based on the sequential parsing of the OCR lines, assuming each row is represented by a block of lines. But to keep it manageable, I'll create a table with the rows that are clearly identifiable.
Given the length, I'll produce a table with all 32 rows as per my speculative list, but note that some fields are uncertain. However, the instruction says not to add or remove substantive words. So I should only use the text present in the OCR. That means I should not invent data. So perhaps the best is to output the OCR text as a series of paragraphs, but that would not be a table.
Wait, the user said: "You are a professional editor specializing in proofreading OCR output... Return the proofread text in standard Markdown... Use Markdown table syntax to reconstruct tabular data." So they want the final output to be a Markdown representation of the document, which includes a table. The OCR output is the raw text. I need to convert that raw text into a proper Markdown table.
Thus, I must interpret the raw text as a table and reconstruct it. I'll do my best to align the data.
Given the complexity, I'll write a script-like approach in my reasoning, but here I'll manually create the table.
Let's look at the raw text again. It might be that the OCR captured the table row by row, but with line breaks within cells. The pattern: each row starts with a name, then the next few lines are the other columns. But the number of lines per row varies because some cells have line breaks.
We can try to detect row boundaries by the presence of a date pattern (e.g., "16th Nov., 1910.") which likely appears in the Date column. Then the next column is Emolument, then Age, then Cause. The Authority column might be before Date.
For Delail Khan, the lines after name: "67.21" (Amount), "16th Nov., 1910." (Date), "210,00" (Emolument), "51" (Age), "Medical Cer-", "tificate." (Cause). So Authority missing. Maybe the Authority column is empty for this row.
For Albert Wilson: "288.00" (Amount), "1st Jan., 1911." (Date), "$64,00" (Emolument). Then next line is "William Robertson," which is next name. So Authority, Age, Cause missing.
For William Robertson: "831,00" (Amount), "10th June, 1911." (Date), "1,662.00" (Emolument). Then next name.
For Amar Singh: "115,00", "11th Aug., 1911.", "300,00". Next name.
For Su Toi: "180.00", "14th Aug., 1911.", "270,00". Next name.
For Hussein Shah: "114.00". Next line "Chu Ping," is a name, so Hussein Shah row has only Amount? Then Date, Emolument, Age, Cause missing? But then "Chu Ping," "81,00" etc. So maybe the table has multiple names with only Amount listed, and then a block of Authority, Date, Emolument, Age, Cause for the group? That seems unlikely.
Look at lines 49-61:
This suggests that the first column (Name) and second column (Amount) are listed for five pensioners, then the third column (Authority) is given once for all? But then Date and Emolument also once? That would be a grouped entry. But then Age and Cause would be per person? The following lines: "* 8 € 2 8", "11", "65", "63", "On Expiration of", "Term of Service.", "75", "쟁", "60", ... These could be Ages and Causes for the five.
"* 8 € 2 8" might be "£8 2s 8d"? Not sure. "11" "65" "63" could be ages. "On Expiration of Term of Service." cause. "75" "쟁" "60" maybe more ages.
This is too ambiguous.
Given the instructions, perhaps the user expects me to output the text as a cleaned-up version of the OCR, preserving the line breaks but fixing words, and not necessarily a perfect table. But they explicitly said "Use Markdown table syntax to reconstruct tabular data." So I must attempt a table.
I'll produce a table with the columns and fill in the data as per the OCR lines in order, using a simple heuristic: each row has 7 cells, and the OCR lines are in cell order. I'll chunk the lines after the header into groups of 7, but the header itself is multiple lines. Let's find where the header ends. The header lines: "Name of Pensioner.", "Amount.", "Amount. Authority.", "Date from which the Pension", "Ins been paid.", "Amount of Emolument when last employed in Public Service.", "Present Age of", "Pensioner.", "Cause of Retirement,", "£", "S. d." That's 11 lines. Then data starts at "Delail Khan,".
If we consider each row has 7 cells, then after the header, the cells are laid out sequentially. But the header has 11 lines, not 7. So maybe the table has 7 columns but the header spans multiple lines.
Another approach: The OCR might have read the table column by column. For example, first all names, then all amounts, then all authorities, etc. But the lines are interleaved.
Given the time constraints, I'll output a Markdown table that includes the data I can parse with high confidence, and for the rest, I'll put the raw lines in a single cell or note.
But the user said "Do not add or remove substantive words." So I should not omit data. I need to include all the text.
Maybe the best is to output the entire OCR text as a Markdown table with one column? No.
I think I need to make a decision. I'll produce a Markdown table with the following columns: Name, Pension Amount, Authority, Date, Emolument, Age, Cause. I'll fill rows by manually matching the data from the OCR lines, using the line numbers as reference. I'll try to be as accurate as possible.
Let's do a systematic parsing by writing a small mental program. But since I can't run code, I'll do it manually for the first few rows and then summarize.
Given the length of the response, I'll produce a table with all the rows I can identify, but I must ensure the output is not excessively long. The user provided a long OCR text. They expect a proofread version. I'll output a Markdown table with the corrected text.
I'll start by creating the header row.
Then for each pensioner, I'll create a row.
I'll use the following data extracted from the OCR lines (with corrections):
324
(L 20)
POLICE PENSIONS.
Name of Pensioner.
Amount.
Amount. Authority.
Date from which the Pension
Ins been paid.
Amount of Emolument when last employed in Public Service.
Present Age of
Pensioner.
Cause of Retirement,
£
S. d.
Delail Khan,
67.21
16th Nov., 1910.
210,00
51
Medical Cer-
tificate.
Albert Wilson,
288.00
1st Jan., 1911.
$64,00
William Robertson,
831,00
10th June, 1911.
1,662.00
Amar Singh,
115,00
11th Aug., 1911.
300,00
Su Toi,
180.00
14th Aug., 1911.
270,00
Hussein Shah,.
114.00
Chu Ping,
81,00
Mahomed Shah,
64.82
Makbool Shah,
65.65
Anokh Singh,.
52.70
Ordinance No. 11 of 1900.
19th Nov.. 1911.
342.00
11
65
63
On Expiration of
Term of Service.
75
争
60
"
1st Jan,
1912.
180.00
53
11
1st Mar., 1912,
210.00
62
Do.
210.00
63
16th Mar., 1912.
186.00
8
Injured on duty.
On Expiration of
Term of service.
50
די
Kesur Singh,
62.95
1st April, 1912.
186.00
64
1
Kadar Bux,.......
52.70
Do.
186.00
46
Lall Singh,
Munga Khan,
Bikh Mahomed,
62.00
1st June, 1912.
186.00
61
Medical Cer- tificate.
¡On Expiration of Term of Service.
Medical Cer- tificate.
60.06
16th July, 1912.
216.00
54
On Expiration of
Term of Service.
$2.70
16th Aug,, 1912.
186.00
49
Medical Cer- tificate.
Kam Kour (widow of
Singh) Rs. 36,
P.C. 769 Jhanda
2 8 0
Granted Special Pension by The Governor in Connell on the death of lier husband who was killed on duty,
19th Ang., 1912,
39
Lau Fuk.....
60.00
1st Sept., 1912.
180,00
50
| Medical Cer-
tificate.
Chan Fan,
Ahmed Deen,
110.19
14th Jan, 1913.
300.00
56
55.80
1st April, 1912.
186,00
པ་
On Expiration of „Term of Service.
Medical Cer- tificate.
Atta Mohamed,
47.12
Do.
186,00
47
**
Nathy Khan, ......
49.60
Do.
186,00
42
19
Henry George Baker,
1,178.57
Li Pui,
60.00
Abdullah,
Mehr Khan,
74.64
55.80
John Lenaghan,
42 0 0
Mahomed Shab,
5.270
Ordinance No. 11 of 1900.
10th April, 1913. 2,016.00
65
On Expiration of
1st July, 1913.
180.00
53
Term of Service.
Medical Certificate.
16th July, 1913.
216.00
48
12
1st Oct., 1913,
186,00
*
48
11
24th Oct,, 1913, £145,0.0 52
29
Gunda Singh,
74.78
Jhande Khun,
136.50
A
William Camerou,
746.76
1st Jan, 1914.
1st Mar., 1914.
1st May, 1914.
23rd May, 1914.
$186.00
55
"
246.00
62
On Expiration of
Terms of Service.
360.00
48
Medical Certificate.
1,632.00
64
fn Expiration of
Term of Service.
Karm Elahi,
95.86
Alfred George Dymond
676.80
1st June, 1914.
6th June, 1914.
246.00
1,632.00
62
29
59
>1
No comments yet.
Private notes are available after approval.