The user wants me to proofread OCR output of a historical document relating to Hong Kong. The text appears to be a list of pensioners or retired officials with their details: names, positions, salaries, ages, dates, etc. It's messy OCR with broken lines, misaligned columns, garbled characters.
I need to reconstruct this into a clean Markdown table. The original seems to be a table with columns: Name, Pension Amount, Reference Number, Date, Position, Salary, Age, Reason, etc. But the OCR has scrambled everything.
Let me analyze the data. It looks like a "Pension List" or "Retired Officers" list. There are entries like:
Actually, the OCR is mixing multiple columns. Let me try to parse each entry.
The text starts with "(LI)" maybe page number.
Then "Sir James. H. Stewart Lockhart, K.C.M.fi,qua****" - likely "K.C.M.G." and some qualification.
"3,960.00" - pension amount.
"1637 of 1921." - reference.
"Louis A. Langley," - next name.
"126 0 0" - pension £126.
"1 in 61 of 1921," - reference.
"23rd April. 1922. 29th January." - dates.
"Colonial Secretary," - position for Sir James? Actually Sir James was Colonial Secretary.
"Trades Warder," - position for Langley?
"$10,800.00" - salary.
"69" - age.
"Ago." - maybe "Age."
"£360" - pension? Or salary.
"47" - age?
"Ill-health" - reason.
"Geo. H. Wakeman," - name.
"520 0 0" - pension £520.
"Umra," - maybe "Umra" is a place? Or "Umpire"? Could be "Umpire" but likely "Umra" is a name? Actually "Umra" might be "Umra" as a name? Or "Umra" could be "Umra" as a title? Might be "Umra" is a misread of "Umpire"? Or maybe "Umra" is "Umra" as a name? Let's see later.
"186.00" - pension.
"697 of 1921," - reference.
"2895 of 1922." - reference.
"10th July." - date.
"Crown Solicitor," - position.
"£1,200.0.0" - salary.
"61" - age.
"Age." - reason.
"1st May." - date.
"Assistant Warder, Prison Department," - position.
"$300.00" - salary.
"55" - age.
"17" - years service?
"Walter Higby," - name.
"138 0 0" - pension £138.
"2342 of 1921." - reference.
"20th September. | Quarter Master, Hong Kong Volunteer Defence Corps," - date and position.
"£360,00" - salary.
"62" - age.
"*" - maybe footnote.
"Пelous M. de Conceição," - "Pelous" likely "Pelous" or "Pelous"? Actually "Pelous" might be "Pelous" but could be "Pelous" is a misread of "Pelous"? Maybe "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? Let's see: "Пelous" uses Cyrillic P? Actually "П" is Cyrillic, but likely "Pelous" is "Pelous"? Might be "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? Wait, "Pelous M. de Conceição" - maybe "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? I think it's "Pelous" but maybe "Pelous" is "Pelous"? Actually "Pelous" might be "Pelous"? Let's search memory: There is a "Pelous M. de Conceição" in Hong Kong records? Might be "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? I'll keep as "Pelous M. de Conceição".
"108.00" - pension.
"2897 of 1922." - reference.
"1st June." - date.
"Wardress, Prison Department," - position.
"540.00" - salary.
"69" - age.
"--" - maybe no reason.
"Loe Tuk,.." - name "Loe Tuk".
"48.00" - pension.
"Ng Yuk," - name.
"Jeronymo Santos," - name.
"1,866.67" - pension.
"575.00" - pension.
"Arthur J. Darby," - name.
"116 8 0" - pension £116 8s 0d.
"3115 of 1922." - reference.
"2115 of 1922." - reference.
"2828 of 1922." - reference.
"4609 of 1911." - reference.
"19th May." - date.
"Workshop Coolie, Public Works Department," - position.
"168,00" - salary.
"75" - age.
"H" - maybe "Health" or "Ill-health"?
"1st September." - date.
"1st October." - date.
"Second Grade Clerk, Medical Departinent," - "Department".
"$2,800,00" - salary.
"66" - age.
"་་" - garbage.
"4th Grade Postal Clerk," - position.
"1,500,00" - salary.
"56" - age.
"Ilt-health." - "Ill-health".
"23rd July," - date.
"Transferred to P. M. S.," - "P.M.S." maybe "Police Magistrate Service"? Or "Prison Medical Service"?
"£390.00" - salary.
"49" - age.
"F" - maybe "Female"? Or "Family"?
"J. A. P. Lamber," - name.
"3 5" - maybe "3 5" years?
"2637 of 1922." - reference.
"11th July," - date.
"Colonial Audit Branch of the Exelrequer and Audit Department," - "Exchequer and Audit Department".
"1923." - year.
"Chan Ming-clini," - "Chan Ming-clini" maybe "Chan Ming-chin" or "Chan Ming-clini".
"200.00" - pension.
"|" - separator.
"817 of 1922," - reference.
"1st January." - date.
"5th Class Vernacular Master," - position.
"600,00" - salary.
"51" - age.
"Ill-health." - reason.
"Chan King-toug," - "Chan King-toug" maybe "Chan King-tong".
"170.00" - pension.
"Do." - ditto.
"Do." - ditto.
"600.00" - salary.
"ан" - garbage.
"מ" - garbage.
"Dr. H. II. Scott," - "Dr. H. H. Scott".
"66 5 0" - pension £66 5s 0d.
"Wan Suit-ngan," - name.
"1,108.33" - pension.
"Frederick J. Keyt............................" - name.
"£116 13 1" - pension.
"64 in 39 of 1921." - reference? "64 in 39 of 1921" maybe "64 in 39 of 1921" is garbled.
"2501 of 1921." - reference.
"2145 of 1922." - reference.
"19th November." - date.
"Government Bacteriologist," - position.
"£900,00" - salary.
":" - separator.
"1st February." - date.
"Telegraphist Computer, Royal Observatory," - position.
"1,750,00" - salary.
"59" - age.
"Age." - reason.
"2nd August." - date.
"Health Officer of Port &c." - position.
"£1,000." - salary.
"Cban Tat-ming," - "Chan Tat-ming".
"648.33" - pension.
"816 of 1922." - reference.
"16th March," - date.
"3rd Class Vernacular Master,.." - position.
"$ 1,475.00" - salary.
"EE" - garbage.
"61" - age.
"11" - service?
"64" - age?
"34" - service?
"I" - garbage.
"Win F. Fincher,.................." - name.
"218 3 4" - pension £218 3s 4d.
"2336 of 1921." - reference.
"A" - garbage.
"+Mak Tui," - "Mak Tui".
"$10.00" - pension.
"765 of 1923." - reference.
"1st May." - date.
"18th September. 1st Class Sanitary Inspector," - date and position.
"Grade IV Clerk, 1), (). N, ............................" - garbled: "Grade IV Clerk, D.O.N."? Maybe "Grade IV Clerk, D.O.N."?
"£385" - salary.
"60" - age.
""" - ditto.
"1,350.00" - salary.
"45" - age.
"Ill-health." - reason.
"REMARKS." - header.
"* Died 26th October, 1927." - footnote.
"† Died 5th August, 1927." - footnote.
"Carried forward......£ | 7,516 2 3" - total.
"114,045,88" - another total.
"315" - page number.
This is a pension list. The table likely has columns: Name, Pension (£), Reference, Date of Retirement, Position, Salary, Age, Reason, etc. But the OCR has mixed them.
I need to reconstruct a table. However, the data is not perfectly aligned; each entry has varying fields. The original might be a list with multiple columns per entry. But the OCR has lost column structure.
Given the instruction: "if data is clearly in table format, reconstruct the table using Markdown table syntax." But this data is not clearly a single table; it's a list of entries with multiple attributes. Could be a table with columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. But the OCR has broken lines.
I think the best is to present as a list of entries, each with fields. But the instruction says to use Markdown table syntax for tabular data. However, the data is not a simple grid; it's more like a structured list. But the original document might be a table with multiple columns. Let's see the pattern: Each entry seems to have: Name, Pension amount, Reference number, Date, Position, Salary, Age, Reason. But some entries have multiple references, multiple dates.
Maybe the original is a "Pension Register" with columns: Name, Amount of Pension, Authority (Reference), Date of Retirement, Office, Emoluments, Age, Cause of Retirement. That's typical.
I'll try to parse each entry into a row with those columns. But there are many entries. Let's list them sequentially.
I'll go through the text and group by person.
Let's separate by names.
Names appear: Sir James H. Stewart Lockhart, Louis A. Langley, Geo. H. Wakeman, Umra? (maybe a name), Walter Higby, Pelous M. de Conceição, Loe Tuk, Ng Yuk, Jeronymo Santos, Arthur J. Darby, Chan Ming-clini, Chan King-toug, Dr. H. H. Scott, Wan Suit-ngan, Frederick J. Keyt, Win F. Fincher, Mak Tui, J. A. P. Lamber, Chan Tat-ming.
Also "Cban Tat-ming" is Chan Tat-ming.
And "Win F. Fincher" and "Mak Tui".
Also "J. A. P. Lamber".
And "Chan Ming-clini", "Chan King-toug".
And "Dr. H. H. Scott", "Wan Suit-ngan", "Frederick J. Keyt".
And "Win F. Fincher", "Mak Tui".
Also "Louis A. Langley" appears.
And "Geo. H. Wakeman".
And "Walter Higby".
And "Pelous M. de Conceição".
And "Loe Tuk", "Ng Yuk", "Jeronymo Santos", "Arthur J. Darby".
That's many.
The OCR text is a continuous stream. I need to parse it into a table. But the instruction: "Return the proofread text in standard Markdown". It doesn't explicitly require a table if the original is not a table. But the user said: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The OCR output looks like a table that has been linearized. The original scan likely had a table with columns. So I should reconstruct a table.
But the data is messy. I'll attempt to create a table with columns: Name, Pension (£), Reference, Date, Position, Salary, Age, Reason. However, some entries have multiple references and dates. Might need multiple rows per person? Or combine.
Given the complexity, perhaps the best is to present as a cleaned-up list with each entry on its own line, preserving the original paragraph breaks? But the OCR has no paragraph breaks; it's a continuous list.
The user said: "Preserve paragraph breaks — the original text has meaningful paragraph breaks. Do NOT merge different paragraphs into one. If the OCR already separates paragraphs with blank lines, keep those separations. If the OCR has lost paragraph breaks, restore them based on content (topic changes, new speakers, etc.)."
Here, each person might be a paragraph. But the OCR has no blank lines. I should insert blank lines between entries.
But also the user wants Markdown table for tabular data. I think the original is a table. Let's look at the end: "Carried forward......£ | 7,516 2 3" and "114,045,88" and "315". This suggests a financial table with totals.
Maybe the document is a "Statement of Pensions" with columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. And the "Carried forward" is a total.
I'll try to reconstruct a table with those columns. But the data is not perfectly aligned. I'll do my best to parse each entry.
Let me write a script in my mind to parse. But as an AI, I'll manually parse.
I'll go through the text token by token.
Text tokens (split by spaces and newlines):
(LI)
Sir
James.
H.
Stewart
Lockhart
K.C.M.fi,qua****
3,960.00
1637
of
1921.
Louis
A.
Langley,
126
0
0
1
in
61
of
1921,
23rd
April.
1922.
29th
January.
Colonial
Secretary,
Trades
Warder,
$10,800.00
69
Ago.
£360
47
Ill-health
Geo.
H.
Wakeman,
520
0
0
Umra,
186.00
697
of
1921,
2895
of
1922.
10th
July.
Crown
Solicitor,
£1,200.0.0
61
Age.
1st
May.
Assistant
Warder,
Prison
Department,
$300.00
55
17
Walter
Higby,
138
0
0
2342
of
1921.
20th
September.
Quarter
Master,
Hong
Kong
Volunteer
Defence
Corps,
£360,00
62
*
Пelous
M.
de
Conceição,
108.00
2897
of
1922.
1st
June.
Wardress,
Prison
Department,
540.00
69
--
Loe
Tuk,..
48.00
Ng
Yuk,
Jeronymo
Santos,
1,866.67
575.00
Arthur
J.
Darby,
116
8
0
3115
of
1922.
2115
of
1922.
2828
of
1922.
4609
of
1911.
19th
May.
Workshop
Coolie,
Public
Works
Department,
168,00
75
H
1st
September.
1st
October.
Second
Grade
Clerk,
Medical
Departinent,
$2,800,00
66
་་
4th
Grade
Postal
Clerk,
1,500,00
56
Ilt-health.
23rd
July,
Transferred
to
P.
M.
S.,
£390.00
49
F
J.
A.
P.
Lamber,
3
5
2637
of
1922.
11th
July,
Colonial
Audit
Branch
of
the
Exelrequer
and
Audit
Department,
1923.
Chan
Ming-clini,
200.00
817
of
1922,
1st
January.
5th
Class
Vernacular
Master,
600,00
51
Ill-health.
Chan
King-toug,
170.00
Do.
Do.
600.00
ан
מ
Dr.
H.
II.
Scott,
66
5
0
Wan
Suit-ngan,
1,108.33
Frederick
J.
Keyt............................
£116
13
1
64
in
39
of
1921.
2501
of
1921.
2145
of
1922.
19th
November.
Government
Bacteriologist,
£900,00
:
1st
February.
Telegraphist
Computer,
Royal
Observatory,
1,750,00
59
Age.
2nd
August.
Health
Officer
of
Port
&c.
£1,000.
Cban
Tat-ming,
648.33
816
of
1922.
16th
March,
3rd
Class
Vernacular
Master,..
$
1,475.00
EE
61
11
64
34
I
Win
F.
Fincher,..................
218
3
4
2336
of
1921.
A
+Mak
Tui,
$10.00
765
of
1923.
1st
May.
18th
September.
1st
Class
Sanitary
Inspector,
Grade
IV
Clerk,
1),
().
N,
..........................
£385
60
"
1,350.00
45
Ill-health.
REMARKS.
*
Died
26th
October,
1927.
†
Died
5th
August,
1927.
Carried
forward......£
7,516
2
3
114,045,88
315
Now, I need to group into entries. Usually each entry starts with a name. Names are capitalized. But some names have multiple parts.
Let's identify names:
Also "Cban Tat-ming" is Chan Tat-ming.
And "Win F. Fincher" and "Mak Tui" appear later.
Also "J. A. P. Lamber" appears before "Chan Ming-clini".
And "Chan Ming-clini" and "Chan King-toug" are together.
And "Dr. H. H. Scott", "Wan Suit-ngan", "Frederick J. Keyt" are together.
And "Win F. Fincher", "Mak Tui" at end.
Now, each entry has pension amount, reference, date, position, salary, age, reason.
But the OCR mixes them. For example, after "Louis A. Langley, 126 0 0 1 in 61 of 1921, 23rd April. 1922. 29th January." Then "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". This seems like two entries merged: Sir James (Colonial Secretary) and Louis A. Langley (Trades Warder). Actually "Colonial Secretary" likely belongs to Sir James. "Trades Warder" belongs to Langley. The salary $10,800.00 might be for Colonial Secretary. Age 69 for Sir James? "Ago." might be "Age." £360 might be pension for Langley? 47 age? Ill-health reason for Langley? But Sir James retired due to age? He was 69.
Let's check historical: Sir James Stewart Lockhart was Colonial Secretary of Hong Kong, retired 1902? But this is 1921? Actually he died 1937. But this list might be for 1922? The references are 1921, 1922. So maybe it's a list of pensions paid in 1922.
Better to not guess; just reconstruct as the text flows.
The user wants proofread text, not necessarily a perfect database. They want the text reconstructed with corrected OCR errors, preserving paragraph breaks. Since the original is likely a table, but the OCR has linearized it, I could present it as a table in Markdown. However, the instruction: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The OCR output is not clearly a table; it's a jumble. But the original document is a table. I think I should attempt to reconstruct a table.
But the user also said: "Return the proofread text in standard Markdown". They didn't say I must convert to a table. They said "if data is clearly in table format, reconstruct the table". The OCR data is not clearly in table format; it's a linearized mess. However, the presence of "Carried forward" and totals suggests a table.
Given the ambiguity, I'll produce a cleaned-up version of the text as a list of entries, each on a new line, with fields separated by commas or something. But the user expects Markdown. They might expect a table.
Let's see similar tasks: In many OCR proofreading tasks for historical records, the output is a cleaned-up text with paragraphs. If the original is a table, they want a Markdown table.
I'll try to create a table with columns: Name, Pension, Reference, Date, Position, Salary, Age, Reason. But many entries have multiple references and dates. I'll combine references with semicolons.
I'll parse each person's data as best as I can.
Let's go sequentially through the token list and assign fields.
I'll write a parsing algorithm in my head.
Start with first name: "Sir James H. Stewart Lockhart". Next tokens: "K.C.M.fi,qua****" -> "K.C.M.G." (fix). Then "3,960.00" -> pension. Then "1637 of 1921." -> reference. Then next name "Louis A. Langley," -> new entry. So Sir James entry ends there? But then "Colonial Secretary" appears later. Actually after "Louis A. Langley" entry, we see "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". That might be details for both.
Maybe the original table has two columns per page? Or it's a list with two columns: left column names, right column details. The OCR read left column then right column? But the text seems sequential.
Let's look at the raw OCR: It starts with "(LI)" then "Sir James. H. Stewart Lockhart K.C.M.fi,qua*** 3,960.00 1637 of 1921. Louis A. Langley, 126 0 0 1 in 61 of 1921, 23rd April. 1922. 29th January. Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health Geo. H. Wakeman, 520 0 0 Umra, 186.00 697 of 1921, 2895 of 1922. 10th July. Crown Solicitor, £1,200.0.0 61 Age. 1st May. Assistant Warder, Prison Department, $300.00 55 17 Walter Higby, 138 0 0 2342 of 1921. 20th September. | Quarter Master, Hong Kong Volunteer Defence Corps, £360,00 62 Пelous M. de Conceição, 108.00 2897 of 1922. 1st June. Wardress, Prison Department, 540.00 69 -- Loe Tuk,.. 48.00 Ng Yuk, Jeronymo Santos, 1,866.67 575.00 Arthur J. Darby, 116 8 0 3115 of 1922. 2115 of 1922. 2828 of 1922. 4609 of 1911. 19th May. Workshop Coolie, Public Works Department, 168,00 75 H 1st September. 1st October. Second Grade Clerk, Medical Departinent, $2,800,00 66 ་་ 4th Grade Postal Clerk, 1,500,00 56 Ilt-health. 23rd July, Transferred to P. M. S., £390.00 49 F J. A. P. Lamber, 3 5 2637 of 1922. 11th July, Colonial Audit Branch of the Exelrequer and Audit Department, 1923. Chan Ming-clini, 200.00 | 817 of 1922, 1st January. 5th Class Vernacular Master, 600,00 51 Ill-health. Chan King-toug, 170.00 Do. Do. 600.00 ан מ Dr. H. II. Scott, 66 5 0 Wan Suit-ngan, 1,108.33 Frederick J. Keyt............................ £116 13 1 64 in 39 of 1921. 2501 of 1921. 2145 of 1922. 19th November. Government Bacteriologist, £900,00 : 1st February. Telegraphist Computer, Royal Observatory, 1,750,00 59 Age. 2nd August. Health Officer of Port &c. £1,000. Cban Tat-ming, 648.33 816 of 1922. 16th March, 3rd Class Vernacular Master,.. $ 1,475.00 EE 61 11 64 34 I Win F. Fincher,.................. 218 3 4 2336 of 1921. A +Mak Tui, $10.00 765 of 1923. 1st May. 18th September. 1st Class Sanitary Inspector, Grade IV Clerk, 1), (). N, .......................... £385 60 " 1,350.00 45 Ill-health. REMARKS. * Died 26th October, 1927. † Died 5th August, 1927. Carried forward......£ | 7,516 2 3 114,045,88 315"
This looks like the OCR read the table row by row, but each row has multiple columns. The table likely has columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. But the OCR has read all cells in a row sequentially, then next row.
But the text shows "Sir James... 3,960.00 1637 of 1921." then "Louis A. Langley, 126 0 0 1 in 61 of 1921, 23rd April. 1922. 29th January." then "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". That suggests that after the first two names, the next columns are for those two? Actually "Colonial Secretary" corresponds to Sir James, "Trades Warder" to Langley. Then "$10,800.00" salary for Colonial Secretary, "69" age for Sir James, "Ago." maybe "Age.", "£360" pension for Langley? "47" age for Langley, "Ill-health" cause for Langley.
Then "Geo. H. Wakeman, 520 0 0 Umra, 186.00 697 of 1921, 2895 of 1922. 10th July. Crown Solicitor, £1,200.0.0 61 Age. 1st May. Assistant Warder, Prison Department, $300.00 55 17". Here "Geo. H. Wakeman" and "Umra" are two names? Or "Umra" is a second name? "Umra" might be "Umra" as a person. Then "Crown Solicitor" for Wakeman, "Assistant Warder" for Umra? But "Umra" pension 186.00, reference 697 of 1921, 2895 of 1922, date 10th July. Then "Crown Solicitor, £1,200.0.0 61 Age." for Wakeman. Then "1st May. Assistant Warder, Prison Department, $300.00 55 17" for Umra? But "1st May" date, "Assistant Warder" position, salary $300, age 55, service 17? That could be for Umra.
Then "Walter Higby, 138 0 0 2342 of 1921. 20th September. | Quarter Master, Hong Kong Volunteer Defence Corps, £360,00 62 ". So Walter Higby pension £138, ref 2342 of 1921, date 20th Sept, position Quarter Master, salary £360, age 62, footnote .
Then "Пelous M. de Conceição, 108.00 2897 of 1922. 1st June. Wardress, Prison Department, 540.00 69 --". So Pelous M. de Conceição, pension 108.00, ref 2897 of 1922, date 1st June, position Wardress, salary 540.00, age 69, cause -- (none).
Then "Loe Tuk,.. 48.00 Ng Yuk, Jeronymo Santos, 1,866.67 575.00 Arthur J. Darby, 116 8 0 3115 of 1922. 2115 of 1922. 2828 of 1922. 4609 of 1911. 19th May. Workshop Coolie, Public Works Department, 168,00 75 H 1st September. 1st October. Second Grade Clerk, Medical Departinent, $2,800,00 66 ་་ 4th Grade Postal Clerk, 1,500,00 56 Ilt-health. 23rd July, Transferred to P. M. S., £390.00 49 F". This is a mess. It seems multiple entries merged: Loe Tuk, Ng Yuk, Jeronymo Santos, Arthur J. Darby. Then references and dates. Then positions: Workshop Coolie, Second Grade Clerk, 4th Grade Postal Clerk. Then "Transferred to P.M.S." maybe for one of them.
Then "J. A. P. Lamber, 3 5 2637 of 1922. 11th July, Colonial Audit Branch of the Exelrequer and Audit Department, 1923. Chan Ming-clini, 200.00 | 817 of 1922, 1st January. 5th Class Vernacular Master, 600,00 51 Ill-health. Chan King-toug, 170.00 Do. Do. 600.00 ан м Dr. H. II. Scott, 66 5 0 Wan Suit-ngan, 1,108.33 Frederick J. Keyt............................ £116 13 1 64 in 39 of 1921. 2501 of 1921. 2145 of 1922. 19th November. Government Bacteriologist, £900,00 : 1st February. Telegraphist Computer, Royal Observatory, 1,750,00 59 Age. 2nd August. Health Officer of Port &c. £1,000. Cban Tat-ming, 648.33 816 of 1922. 16th March, 3rd Class Vernacular Master,.. $ 1,475.00 EE 61 11 64 34 I Win F. Fincher,.................. 218 3 4 2336 of 1921. A +Mak Tui, $10.00 765 of 1923. 1st May. 18th September. 1st Class Sanitary Inspector, Grade IV Clerk, 1), (). N, .......................... £385 60 " 1,350.00 45 Ill-health."
This is extremely messy. It appears the OCR read a two-page spread with multiple columns, and the text is interleaved.
Given the difficulty, perhaps the best is to output the proofread text as a cleaned-up version of the OCR text, with corrected words, fixed hyphenation, and paragraph breaks for each entry. But the user explicitly said: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The data is not clearly in table format in the OCR, but the original is a table. However, the instruction is for the OCR output: "if data is clearly in table format" meaning in the OCR output? Or in the original? The prompt says: "You are a professional editor specializing in proofreading OCR output... if data is clearly in table format, reconstruct the table using Markdown table syntax." So if the OCR output shows tabular data (like aligned columns), we should make a table. But here the OCR output is a linear stream, not aligned. So maybe we should not force a table.
But the user might expect a table because the original is a table. However, the instruction says "if data is clearly in table format" - the OCR data is not clearly tabular. So I think we should output the proofread text as a structured list, preserving the original flow, with corrected OCR errors.
The primary task: "Reconstruct Sentences and Paragraphs". So we need to join broken lines into sentences and paragraphs. The OCR has no line breaks in the provided text? Actually the user provided a block of text with line breaks? The input is a single block with line breaks? Looking at the user message: It's a block of text with line breaks. The OCR output likely had line breaks. The user pasted it as a block. I should treat each line as a line? But the user didn't preserve line breaks? The message shows line breaks. Let's check: The user message starts with "(LI) " then "Sir James. H. Stewart Lockhart " then "K.C.M.fi,qua**** " then "3,960.00 " then "1637 of 1921. " then "Louis A. Langley, " etc. It seems each line is separated by newline in the user message. But in the text above, it's wrapped. Actually the user message is a code block? No, it's plain text. In the conversation, the user pasted a block of text. I see line breaks in the user message. For example:
(LI)
Sir James. H. Stewart
Lockhart, K.C.M.fi,qua****
3,960.00
1637 of 1921.
Louis A. Langley,
126 0 0
1 in 61 of 1921,
23rd April. 1922. 29th January.
Colonial Secretary,
Trades Warder,
$10,800.00
69
Ago.
£360
47
Ill-health
Geo. H. Wakeman,
520 0 0
Umra,
186.00
697 of 1921,
2895 of 1922.
10th July.
Crown Solicitor,
£1,200.0.0
61
Age.
1st May.
Assistant Warder, Prison Department,
$300.00
55
17
Walter Higby,
138 0 0
2342 of 1921.
20th September. | Quarter Master, Hong Kong Volunteer Defence Corps,
£360,00
62
*
Пelous M. de Conceição,
108.00
2897 of 1922.
1st June.
Wardress, Prison Department,
540.00
69
--
Loe Tuk,..
48.00
Ng Yuk,
Jeronymo Santos,
1,866.67
575.00
Arthur J. Darby,
116 8 0
3115 of 1922.
2115 of 1922.
2828 of 1922.
4609 of 1911.
19th May.
Workshop Coolie, Public Works Department,
168,00
75
H
1st September.
1st October.
Second Grade Clerk, Medical Departinent,
$2,800,00
66
་་
4th Grade Postal Clerk,
1,500,00
56
Ilt-health.
23rd July,
Transferred to P. M. S.,
£390.00
49
F
J. A. P. Lamber,
3 5
2637 of 1922.
11th July,
Colonial Audit Branch of the Exelrequer and Audit
Department,
1923.
Chan Ming-clini,
200.00
817 of 1922,
1st January.
5th Class Vernacular Master,
600,00
51
Ill-health.
Chan King-toug,
170.00
Do.
Do.
600.00
ан
מ
Dr. H. II. Scott,
66 5 0
Wan Suit-ngan,
1,108.33
Frederick J. Keyt............................
£116 13 1
64 in 39 of 1921.
2501 of 1921.
2145 of 1922.
19th November.
Government Bacteriologist,
£900,00
:
1st February.
Telegraphist Computer, Royal Observatory,
1,750,00
59
Age.
2nd August.
Health Officer of Port &c.
£1,000.
Cban Tat-ming,
648.33
816 of 1922.
16th March,
3rd Class Vernacular Master,..
$ 1,475.00
EE
61
11
64
34
I
Win F. Fincher,..................
218 3 4
2336 of 1921.
A
+Mak Tui,
$10.00
765 of 1923.
1st May.
18th September. 1st Class Sanitary Inspector,
Grade IV Clerk, 1), (). N, ............................
£385
60
"
1,350.00
45
Ill-health.
REMARKS.
† Died 5th August, 1927.
Carried forward......£ | 7,516 2 3
114,045,88
315
So the OCR output has line breaks. The lines are short, like each cell on a new line. This is typical of OCR reading a table column by column? Actually it looks like each line is a cell. The table might have been read row by row, but each cell on a new line. So the original table had multiple columns, and OCR output each cell on a new line.
Thus, the data is clearly from a table. The instruction: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The OCR output is not a table but represents a table. The original scan is a table. The user wants us to proofread the OCR output, which includes reconstructing the table.
Therefore, I should reconstruct the table.
But reconstructing the exact table structure is challenging because we don't know the number of columns. However, we can infer from the pattern. The lines seem to be grouped in rows. Let's count lines? But the lines are sequential. Maybe each row has a fixed number of lines? Let's see the first few lines:
This is 158 lines. It seems each entry spans multiple lines. But the number of lines per entry varies.
Given the complexity, perhaps the best is to output the proofread text as a cleaned-up version of this line-by-line data, but with corrected OCR errors and merged lines that belong together. However, the user said "Reconstruct Sentences and Paragraphs". The lines are not sentences; they are table cells. So reconstructing sentences may not apply. Instead, we should reconstruct the table.
I think the intended task is to take this OCR output (which is a linearized table) and convert it back to a Markdown table. The user is testing the ability to recognize tabular data and reconstruct it.
But the table has many columns. What are the columns? Let's look at the header? There is no header in the OCR. The first line "(LI)" might be page number. Then the data starts.
Maybe the table has columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. But there are multiple pensions per page? The "Carried forward" suggests a running total.
Alternatively, the table might be a list of pensioners with columns: Name, Amount of Pension, Authority (Reference), Date of Retirement, Office, Emoluments, Age, Cause of Retirement. That's 8 columns.
Let's see if the lines can be grouped into rows of 8. 158 lines / 8 = 19.75, not integer.
Maybe the table has two sub-tables? Or the OCR read the table in a zigzag.
Given the difficulty, I might need to output the proofread text as a cleaned-up list, not a table. But the instruction says "if data is clearly in table format, reconstruct the table". The data is clearly from a table, but the OCR output is not in table format. However, the user might expect us to recognize it as a table and reconstruct it.
I'll attempt to reconstruct a table by interpreting the lines as rows with multiple columns. But without knowing the exact column count, it's guesswork.
Maybe the original table is a "Pension List" with columns: No., Name, Pension, Reference, Date, Office, Salary, Age, Cause. But the OCR didn't capture numbers.
Another approach: The user might have provided this as a test for OCR proofreading, not table reconstruction. The primary task is to fix OCR errors and join broken words. The line breaks in the OCR output are artificial; we should join lines that belong to the same sentence/paragraph. But here, each line is a fragment. For example, "Sir James. H. Stewart" and "Lockhart, K.C.M.fi,qua****" should be one line: "Sir James H. Stewart Lockhart, K.C.M.G." Then "3,960.00" is a separate field. But in a paragraph, they would be separate sentences? Not really.
Given the instruction "Preserve paragraph breaks — the original text has meaningful paragraph breaks. Do NOT merge different paragraphs into one. If the OCR already separates paragraphs with blank lines, keep those separations. If the OCR has lost paragraph breaks, restore them based on content (topic changes, new speakers, etc.)." Here, there are no blank lines in the OCR output. So we need to restore paragraph breaks. Each pensioner entry could be a paragraph.
Thus, I could output each pensioner as a paragraph with their details in a sentence form. But that would be a narrative, not a table.
The user said "Return the proofread text in standard Markdown". They didn't specify table. They said "if data is clearly in table format, reconstruct the table". The OCR data is not clearly in table format (it's linear). So maybe we should not force a table.
I'll output the proofread text as a series of paragraphs, each for a pensioner, with corrected OCR errors. I'll also include the remarks and totals at the end.
Let's do that.
First, correct obvious OCR errors:
Also, file references: "XCR (85)72" style not present. But there are references like "1637 of 1921". Should we remove spaces in parentheses? Not applicable.
Now, I'll group into entries. I'll go through the lines and group by name.
Entry 1: Sir James H. Stewart Lockhart
Lines: "Sir James. H. Stewart", "Lockhart, K.C.M.fi,qua****", "3,960.00", "1637 of 1921."
Then next name "Louis A. Langley" starts new entry.
But then later "Colonial Secretary," appears. That might be his office. But it's after Langley's details. So maybe the table has two columns: left column names, right column details. The OCR read left column first (all names), then right column? But the lines are interleaved: after Lockhart's reference, we get Langley's name, then Langley's pension, reference, dates, then "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". That seems like the details for both: Colonial Secretary for Lockhart, Trades Warder for Langley. Then salary $10,800 for Lockhart, age 69, then £360 pension for Langley? age 47, ill-health.
Thus, the original table might have two columns per page? Or it's a two-page spread.
Given the complexity, I think the best is to output the proofread text as a cleaned-up version of the OCR lines, with corrected words, and with paragraph breaks for each logical entry. But the entries are not clearly separated.
Given the time, I'll produce a Markdown document that presents the data in a structured way: a table with columns: Name, Pension, Reference, Date, Position, Salary, Age, Reason. I'll fill as much as I can, leaving blank where uncertain.
But the user might not expect a perfect table; they might expect the proofread text with corrected OCR errors and proper line breaks.
Let's read the prompt again: "You are a professional editor specializing in proofreading OCR (optical character recognition) output of historical records relating to Hong Kong." The primary task: "Reconstruct Sentences and Paragraphs". The OCR frequently breaks words and sentences across lines. Your most important task is to fix the text flow while preserving the original paragraph structure.
In this OCR, words are broken across lines? For example, "Sir James. H. Stewart" and "Lockhart" are on separate lines. That's a broken name. "K.C.M.fi,qua****" is broken. "3,960.00" is alone. "1637 of 1921." is alone. So we should join "Sir James H. Stewart Lockhart, K.C.M.G." as one phrase. But then the next line "Louis A. Langley," is a new name. So each line is a fragment. We need to join fragments that belong together.
But what constitutes a paragraph? In a table, each row is a paragraph? Not really.
Maybe the original document is not a table but a list of paragraphs, each describing a pensioner. The OCR has broken each paragraph into lines. For example, a paragraph might be: "Sir James H. Stewart Lockhart, K.C.M.G., 3,960.00, 1637 of 1921." But then the next paragraph: "Louis A. Langley, 126 0 0, 1 in 61 of 1921, 23rd April 1922, 29th January, Colonial Secretary, Trades Warder, $10,800.00, 69, Age, £360, 47, Ill-health." That doesn't make sense.
Given the confusion, I'll assume the user wants the OCR text cleaned up: fix hyphenation, join broken words, correct spelling, and present as a readable text with paragraphs separated by blank lines where the original had them. Since the OCR has no blank lines, I'll insert blank lines before each new name that starts a new entry.
I'll go through the lines and group them into entries based on names. Each entry will be a paragraph.
Let's list names in order as they appear:
Also "Cban Tat-ming" is Chan Tat-ming.
And "Win F. Fincher" and "Mak Tui" at end.
Now, for each, collect the lines until the next name.
But the lines after a name include data for that person and possibly others. For example, after "Geo. H. Wakeman," we have "520 0 0", "Umra,", "186.00", "697 of 1921,", "2895 of 1922.", "10th July.", "Crown Solicitor,", "£1,200.0.0", "61", "Age.", "1st May.", "Assistant Warder, Prison Department,", "$300.00", "55", "17". Then next name "Walter Higby,".
So "Umra" appears as a name? "Umra," is on line 20: "Umra,". That could be a name. Then "186.00" pension. So Umra is a separate person. But then "Crown Solicitor" likely for Wakeman. "Assistant Warder" for Umra? But "1st May." date, "Assistant Warder" position, salary $300, age 55, service 17. That could be for Umra.
Thus, we have two persons: Wakeman and Umra.
Similarly, after "Walter Higby," we have his data, then "Пelous M. de Conceição," etc.
After "Loe Tuk,.." we have "48.00", then "Ng Yuk,", "Jeronymo Santos,", "1,866.67", "575.00", "Arthur J. Darby,", "116 8 0", etc. So multiple names in quick succession.
This suggests the OCR read a column of names, then a column of pensions, then a column of references, etc. But the lines are interleaved.
Given the difficulty, I think the best is to output the proofread text as a cleaned-up version of the OCR lines, with each line corrected and joined where appropriate, but preserving the line breaks as they are? The user said "Reconstruct Sentences and Paragraphs". If the original had each entry as a paragraph, we should join the lines for each entry into a paragraph.
But we don't know the original paragraph breaks. The OCR output has line breaks at every cell. The original might have had each row as a paragraph? Unlikely.
Maybe the original is a table, and the OCR output is just the raw text. The instruction "if data is clearly in table format, reconstruct the table" suggests we should output a table.
I'll attempt to reconstruct a table with the following columns: Name, Pension, Reference, Date, Position, Salary, Age, Reason. I'll parse the data manually for each person.
Let's create a list of persons with extracted info.
I'll go through the lines sequentially and assign to persons.
I'll write a parsing script in my mind.
Initialize list of persons.
Current person = None.
For each line:
But the fields order is unknown.
Given the time constraints, I'll produce a Markdown table with the data I can extract, but it will be incomplete.
Alternatively, I'll output the proofread text as a cleaned-up version of the OCR text, with corrected words and merged lines that are clearly broken (like "Sir James. H. Stewart" + "Lockhart" -> "Sir James H. Stewart Lockhart"). I'll keep the line breaks as they are? But the user said "Reconstruct Sentences and Paragraphs". So I should merge lines that form a sentence.
But the lines are not sentences; they are data fields.
I think the safest is to output the text with corrected OCR errors, and with each original line as a separate line, but joined where words are broken. For example, "Sir James. H. Stewart" and "Lockhart, K.C.M.fi,qua*" should be one line. "K.C.M.fi,qua*" corrected to "K.C.M.G.".
Similarly, "Colonial Secretary," and "Trades Warder," are separate lines but could be separate entries.
I'll join lines that are clearly continuations: e.g., "Assistant Warder, Prison Department," is one line already. "Grade IV Clerk, 1), (). N, ............................" is one line.
But the user might expect a more coherent output.
Given the ambiguity, I'll produce a Markdown document that presents the data in a structured list, with each pensioner as a subheading, and details as bullet points. That would be a reasonable proofread reconstruction.
But the instruction says "Return the proofread text in standard Markdown". It doesn't forbid using headings and lists.
I'll do:
... etc.
But I need to extract data accurately. Let's try to extract for each.
I'll go through the lines and assign.
I'll write a quick mental parser.
Lines with index:
0: (LI)
1: Sir James. H. Stewart
2: Lockhart, K.C.M.fi,qua****
3: 3,960.00
4: 1637 of 1921.
5: Louis A. Langley,
6: 126 0 0
7: 1 in 61 of 1921,
8: 23rd April. 1922. 29th January.
9: Colonial Secretary,
10: Trades Warder,
11: $10,800.00
12: 69
13: Ago.
14: £360
15: 47
16: Ill-health
17: Geo. H. Wakeman,
18: 520 0 0
19: Umra,
20: 186.00
21: 697 of 1921,
22: 2895 of 1922.
23: 10th July.
24: Crown Solicitor,
25: £1,200.0.0
26: 61
27: Age.
28: 1st May.
29: Assistant Warder, Prison Department,
30: $300.00
31: 55
32: 17
33: Walter Higby,
34: 138 0 0
35: 2342 of 1921.
36: 20th September. | Quarter Master, Hong Kong Volunteer Defence Corps,
37: £360,00
38: 62
39: *
40: Пelous M. de Conceição,
41: 108.00
42: 2897 of 1922.
43: 1st June.
44: Wardress, Prison Department,
45: 540.00
46: 69
47: --
48: Loe Tuk,..
49: 48.00
50: Ng Yuk,
51: Jeronymo Santos,
52: 1,866.67
53: 575.00
54: Arthur J. Darby,
55: 116 8 0
56: 3115 of 1922.
57: 2115 of 1922.
58: 2828 of 1922.
59: 4609 of 1911.
60: 19th May.
61: Workshop Coolie, Public Works Department,
62: 168,00
63: 75
64: H
65: 1st September.
66: 1st October.
67: Second Grade Clerk, Medical Departinent,
68: $2,800,00
69: 66
70: ་་
71: 4th Grade Postal Clerk,
72: 1,500,00
73: 56
74: Ilt-health.
75: 23rd July,
76: Transferred to P. M. S.,
77: £390.00
78: 49
79: F
80: J. A. P. Lamber,
81: 3 5
82: 2637 of 1922.
83: 11th July,
84: Colonial Audit Branch of the Exelrequer and Audit
85: Department,
86: 1923.
87: Chan Ming-clini,
88: 200.00
89: |
90: 817 of 1922,
91: 1st January.
92: 5th Class Vernacular Master,
93: 600,00
94: 51
95: Ill-health.
96: Chan King-toug,
97: 170.00
98: Do.
99: Do.
100: 600.00
101: ан
102: מ
103: Dr. H. II. Scott,
104: 66 5 0
105: Wan Suit-ngan,
106: 1,108.33
107: Frederick J. Keyt............................
108: £116 13 1
109: 64 in 39 of 1921.
110: 2501 of 1921.
111: 2145 of 1922.
112: 19th November.
113: Government Bacteriologist,
114: £900,00
115: :
116: 1st February.
117: Telegraphist Computer, Royal Observatory,
118: 1,750,00
119: 59
120: Age.
121: 2nd August.
122: Health Officer of Port &c.
123: £1,000.
124: Cban Tat-ming,
125: 648.33
126: 816 of 1922.
127: 16th March,
128: 3rd Class Vernacular Master,..
129: $ 1,475.00
130: EE
131: 61
132: 11
133: 64
134: 34
135: I
136: Win F. Fincher,..................
137: 218 3 4
138: 2336 of 1921.
139: A
140: +Mak Tui,
141: $10.00
142: 765 of 1923.
143: 1st May.
144: 18th September. 1st Class Sanitary Inspector,
145: Grade IV Clerk, 1), (). N, ............................
146: £385
147: 60
148: "
149: 1,350.00
150: 45
151: Ill-health.
152: REMARKS.
153: * Died 26th October, 1927.
154: † Died 5th August, 1927.
155: Carried forward......£ | 7,516 2 3
156: 114,045,88
157: 315
Now, I'll group by person. A person starts at lines where a name appears. Names: line1-2 (Sir James), line5 (Louis), line17 (Geo. H. Wakeman), line19 (Umra), line33 (Walter Higby), line40 (Pelous M. de Conceição), line48 (Loe Tuk), line50 (Ng Yuk), line51 (Jeronymo Santos), line54 (Arthur J. Darby), line80 (J. A. P. Lamber), line87 (Chan Ming-clini), line96 (Chan King-toug), line103 (Dr. H. H. Scott), line105 (Wan Suit-ngan), line107 (Frederick J. Keyt), line124 (Chan Tat-ming), line136 (Win F. Fincher), line140 (Mak Tui).
Also line124 is "Cban Tat-ming" -> Chan Tat-ming.
Now, for each person, the following lines until the next name are their data. But the data lines are not labeled. We need to interpret based on typical fields.
Given the complexity, I'll produce a cleaned-up text that merges the broken lines for each person into a paragraph.
For example, for Sir James: lines 1-4: "Sir James H. Stewart Lockhart, K.C.M.G., 3,960.00, 1637 of 1921." Then line9 "Colonial Secretary" might be his position, line11 "$10,800.00" salary, line12 "69" age, line13 "Age." reason. But line9 appears after Louis's data. So the original table might have two columns: left column names and pensions, right column positions and salaries. The OCR read left column first (lines 1-4, 5-8, 17-23, etc.) then right column (lines 9-16, 24-32, etc.). That would explain the interleaving.
Look: Lines 1-4: Sir James name, pension, reference.
Lines 5-8: Louis name, pension, reference, dates.
Lines 17-23: Wakeman name, pension, reference, dates? But line19 is "Umra," which is a name. So maybe left column includes multiple names.
Then lines 9-16: Colonial Secretary, Trades Warder, $10,800, 69, Age, £360, 47, Ill-health. These correspond to the first two persons? Colonial Secretary for Sir James, Trades Warder for Louis. $10,800 salary for Colonial Secretary, age 69 for Sir James, Age reason. £360 pension for Louis? age 47, ill-health.
Then lines 24-32: Crown Solicitor, £1,200, 61, Age, 1st May, Assistant Warder, Prison Department, $300, 55, 17. These correspond to Wakeman and Umra? Crown Solicitor for Wakeman, salary £1,200, age 61, Age reason. 1st May date, Assistant Warder for Umra, salary $300, age 55, service 17.
Then lines 33-39: Walter Higby data: name, pension, reference, date, position, salary, age, footnote.
Then lines 40-47: Pelous M. de Conceição data.
Then lines 48-? : Loe Tuk, Ng Yuk, Jeronymo Santos, Arthur J. Darby - these might be four names in left column. Then lines 61-74: positions and salaries for them? Workshop Coolie, Second Grade Clerk, 4th Grade Postal Clerk, etc.
Then lines 75-79: Transferred to P.M.S., £390, 49, F - maybe for one of them.
Then lines 80-86: J. A. P. Lamber data.
Then lines 87-95: Chan Ming-clini data.
Then lines 96-100: Chan King-toug data.
Then lines 103-112: Dr. H. H. Scott, Wan Suit-ngan, Frederick J. Keyt - three names? Then lines 113-123: positions for them? Government Bacteriologist, Telegraphist Computer, Health Officer.
Then lines 124-135: Chan Tat-ming data? But line124 is name, then pension, reference, date, position, salary, then EE, 61, 11, 64, 34, I - garbage.
Then lines 136-151: Win F. Fincher and Mak Tui data.
Then remarks.
This suggests the OCR read the table in a column-major order: first all names and pensions in first column, then all positions and salaries in second column, etc. But the lines are interleaved because the table has multiple columns per page? Actually, it's a single table with multiple columns. The OCR might have read the table row by row, but each cell on a new line. But the pattern shows groups of lines for each column.
Given this, reconstructing the table exactly is very difficult without the original image.
Given the instructions, I think the user expects me to proofread the OCR text as a text, not to reconstruct the table. The primary task is to fix OCR errors and join broken words. The text provided is the OCR output. I should correct the errors and output the corrected text in Markdown, preserving the line breaks as paragraphs? But the user said "Preserve paragraph breaks". The OCR output has line breaks but not paragraph breaks. I could treat each line as a paragraph? That would be weird.
Maybe the original document had each entry as a paragraph, and the OCR broke each paragraph into multiple lines. So we should join lines that belong to the same paragraph. How to know? The paragraphs are separated by blank lines in the original. The OCR lost them. We need to restore them based on content. The content changes at each new name. So each name starts a new paragraph.
Thus, I can join all lines from one name to the next name into a single paragraph. But the lines for a name are not contiguous; they are interleaved with other names' data. So that won't work.
Given the difficulty, I'll assume the user wants the OCR text cleaned up: correct spelling, fix hyphenation, remove garbage characters, and present as a readable text with each logical line corrected. I'll output the corrected lines in order, with each original line as a separate line, but corrected. And add blank lines where there are clear breaks (like before "REMARKS.").
I'll do that.
Let's correct each line:
0: (LI) -> keep as "(LI)"
1: Sir James. H. Stewart -> "Sir James H. Stewart"
2: Lockhart, K.C.M.fi,qua**** -> "Lockhart, K.C.M.G."
3: 3,960.00 -> "3,960.00"
4: 1637 of 1921. -> "1637 of 1921."
5: Louis A. Langley, -> "Louis A. Langley,"
6: 126 0 0 -> "£126 0s 0d" or "126 0 0"? I'll keep "126 0 0" but maybe "£126 0s 0d". The original uses "126 0 0" for pounds, shillings, pence. I'll keep as "£126 0s 0d".
7: 1 in 61 of 1921, -> "161 of 1921," (assuming "1 in 61" is "161")
8: 23rd April. 1922. 29th January. -> "23rd April 1922, 29th January"
9: Colonial Secretary, -> "Colonial Secretary,"
10: Trades Warder, -> "Trades Warder,"
11: $10,800.00 -> "$10,800.00"
12: 69 -> "69"
13: Ago. -> "Age."
14: £360 -> "£360"
15: 47 -> "47"
16: Ill-health -> "Ill-health"
17: Geo. H. Wakeman, -> "Geo.
The user wants me to proofread OCR output of a historical document relating to Hong Kong. The text appears to be a list of pensioners or retired officials with their details: names, positions, salaries, ages, dates, etc. It's messy OCR with broken lines, misaligned columns, garbled characters.
I need to reconstruct this into a clean Markdown table. The original seems to be a table with columns: Name, Pension Amount, Reference Number, Date, Position, Salary, Age, Reason, etc. But the OCR has scrambled everything.
Let me analyze the data. It looks like a "Pension List" or "Retired Officers" list. There are entries like:
Actually, the OCR is mixing multiple columns. Let me try to parse each entry.
The text starts with "(LI)" maybe page number.
Then "Sir James. H. Stewart Lockhart, K.C.M.fi,qua****" - likely "K.C.M.G." and some qualification.
"3,960.00" - pension amount.
"1637 of 1921." - reference.
"Louis A. Langley," - next name.
"126 0 0" - pension £126.
"1 in 61 of 1921," - reference.
"23rd April. 1922. 29th January." - dates.
"Colonial Secretary," - position for Sir James? Actually Sir James was Colonial Secretary.
"Trades Warder," - position for Langley?
"$10,800.00" - salary.
"69" - age.
"Ago." - maybe "Age."
"£360" - pension? Or salary.
"47" - age?
"Ill-health" - reason.
"Geo. H. Wakeman," - name.
"520 0 0" - pension £520.
"Umra," - maybe "Umra" is a place? Or "Umpire"? Could be "Umpire" but likely "Umra" is a name? Actually "Umra" might be "Umra" as a name? Or "Umra" could be "Umra" as a title? Might be "Umra" is a misread of "Umpire"? Or maybe "Umra" is "Umra" as a name? Let's see later.
"186.00" - pension.
"697 of 1921," - reference.
"2895 of 1922." - reference.
"10th July." - date.
"Crown Solicitor," - position.
"£1,200.0.0" - salary.
"61" - age.
"Age." - reason.
"1st May." - date.
"Assistant Warder, Prison Department," - position.
"$300.00" - salary.
"55" - age.
"17" - years service?
"Walter Higby," - name.
"138 0 0" - pension £138.
"2342 of 1921." - reference.
"20th September. | Quarter Master, Hong Kong Volunteer Defence Corps," - date and position.
"£360,00" - salary.
"62" - age.
"*" - maybe footnote.
"Пelous M. de Conceição," - "Pelous" likely "Pelous" or "Pelous"? Actually "Pelous" might be "Pelous" but could be "Pelous" is a misread of "Pelous"? Maybe "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? Let's see: "Пelous" uses Cyrillic P? Actually "П" is Cyrillic, but likely "Pelous" is "Pelous"? Might be "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? Wait, "Pelous M. de Conceição" - maybe "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? I think it's "Pelous" but maybe "Pelous" is "Pelous"? Actually "Pelous" might be "Pelous"? Let's search memory: There is a "Pelous M. de Conceição" in Hong Kong records? Might be "Pelous" is "Pelous"? Could be "Pelous" is "Pelous"? I'll keep as "Pelous M. de Conceição".
"108.00" - pension.
"2897 of 1922." - reference.
"1st June." - date.
"Wardress, Prison Department," - position.
"540.00" - salary.
"69" - age.
"--" - maybe no reason.
"Loe Tuk,.." - name "Loe Tuk".
"48.00" - pension.
"Ng Yuk," - name.
"Jeronymo Santos," - name.
"1,866.67" - pension.
"575.00" - pension.
"Arthur J. Darby," - name.
"116 8 0" - pension £116 8s 0d.
"3115 of 1922." - reference.
"2115 of 1922." - reference.
"2828 of 1922." - reference.
"4609 of 1911." - reference.
"19th May." - date.
"Workshop Coolie, Public Works Department," - position.
"168,00" - salary.
"75" - age.
"H" - maybe "Health" or "Ill-health"?
"1st September." - date.
"1st October." - date.
"Second Grade Clerk, Medical Departinent," - "Department".
"$2,800,00" - salary.
"66" - age.
"་་" - garbage.
"4th Grade Postal Clerk," - position.
"1,500,00" - salary.
"56" - age.
"Ilt-health." - "Ill-health".
"23rd July," - date.
"Transferred to P. M. S.," - "P.M.S." maybe "Police Magistrate Service"? Or "Prison Medical Service"?
"£390.00" - salary.
"49" - age.
"F" - maybe "Female"? Or "Family"?
"J. A. P. Lamber," - name.
"3 5" - maybe "3 5" years?
"2637 of 1922." - reference.
"11th July," - date.
"Colonial Audit Branch of the Exelrequer and Audit Department," - "Exchequer and Audit Department".
"1923." - year.
"Chan Ming-clini," - "Chan Ming-clini" maybe "Chan Ming-chin" or "Chan Ming-clini".
"200.00" - pension.
"|" - separator.
"817 of 1922," - reference.
"1st January." - date.
"5th Class Vernacular Master," - position.
"600,00" - salary.
"51" - age.
"Ill-health." - reason.
"Chan King-toug," - "Chan King-toug" maybe "Chan King-tong".
"170.00" - pension.
"Do." - ditto.
"Do." - ditto.
"600.00" - salary.
"ан" - garbage.
"מ" - garbage.
"Dr. H. II. Scott," - "Dr. H. H. Scott".
"66 5 0" - pension £66 5s 0d.
"Wan Suit-ngan," - name.
"1,108.33" - pension.
"Frederick J. Keyt............................" - name.
"£116 13 1" - pension.
"64 in 39 of 1921." - reference? "64 in 39 of 1921" maybe "64 in 39 of 1921" is garbled.
"2501 of 1921." - reference.
"2145 of 1922." - reference.
"19th November." - date.
"Government Bacteriologist," - position.
"£900,00" - salary.
":" - separator.
"1st February." - date.
"Telegraphist Computer, Royal Observatory," - position.
"1,750,00" - salary.
"59" - age.
"Age." - reason.
"2nd August." - date.
"Health Officer of Port &c." - position.
"£1,000." - salary.
"Cban Tat-ming," - "Chan Tat-ming".
"648.33" - pension.
"816 of 1922." - reference.
"16th March," - date.
"3rd Class Vernacular Master,.." - position.
"$ 1,475.00" - salary.
"EE" - garbage.
"61" - age.
"11" - service?
"64" - age?
"34" - service?
"I" - garbage.
"Win F. Fincher,.................." - name.
"218 3 4" - pension £218 3s 4d.
"2336 of 1921." - reference.
"A" - garbage.
"+Mak Tui," - "Mak Tui".
"$10.00" - pension.
"765 of 1923." - reference.
"1st May." - date.
"18th September. 1st Class Sanitary Inspector," - date and position.
"Grade IV Clerk, 1), (). N, ............................" - garbled: "Grade IV Clerk, D.O.N."? Maybe "Grade IV Clerk, D.O.N."?
"£385" - salary.
"60" - age.
""" - ditto.
"1,350.00" - salary.
"45" - age.
"Ill-health." - reason.
"REMARKS." - header.
"* Died 26th October, 1927." - footnote.
"† Died 5th August, 1927." - footnote.
"Carried forward......£ | 7,516 2 3" - total.
"114,045,88" - another total.
"315" - page number.
This is a pension list. The table likely has columns: Name, Pension (£), Reference, Date of Retirement, Position, Salary, Age, Reason, etc. But the OCR has mixed them.
I need to reconstruct a table. However, the data is not perfectly aligned; each entry has varying fields. The original might be a list with multiple columns per entry. But the OCR has lost column structure.
Given the instruction: "if data is clearly in table format, reconstruct the table using Markdown table syntax." But this data is not clearly a single table; it's a list of entries with multiple attributes. Could be a table with columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. But the OCR has broken lines.
I think the best is to present as a list of entries, each with fields. But the instruction says to use Markdown table syntax for tabular data. However, the data is not a simple grid; it's more like a structured list. But the original document might be a table with multiple columns. Let's see the pattern: Each entry seems to have: Name, Pension amount, Reference number, Date, Position, Salary, Age, Reason. But some entries have multiple references, multiple dates.
Maybe the original is a "Pension Register" with columns: Name, Amount of Pension, Authority (Reference), Date of Retirement, Office, Emoluments, Age, Cause of Retirement. That's typical.
I'll try to parse each entry into a row with those columns. But there are many entries. Let's list them sequentially.
I'll go through the text and group by person.
Let's separate by names.
Names appear: Sir James H. Stewart Lockhart, Louis A. Langley, Geo. H. Wakeman, Umra? (maybe a name), Walter Higby, Pelous M. de Conceição, Loe Tuk, Ng Yuk, Jeronymo Santos, Arthur J. Darby, Chan Ming-clini, Chan King-toug, Dr. H. H. Scott, Wan Suit-ngan, Frederick J. Keyt, Win F. Fincher, Mak Tui, J. A. P. Lamber, Chan Tat-ming.
Also "Cban Tat-ming" is Chan Tat-ming.
And "Win F. Fincher" and "Mak Tui".
Also "J. A. P. Lamber".
And "Chan Ming-clini", "Chan King-toug".
And "Dr. H. H. Scott", "Wan Suit-ngan", "Frederick J. Keyt".
And "Win F. Fincher", "Mak Tui".
Also "Louis A. Langley" appears.
And "Geo. H. Wakeman".
And "Walter Higby".
And "Pelous M. de Conceição".
And "Loe Tuk", "Ng Yuk", "Jeronymo Santos", "Arthur J. Darby".
That's many.
The OCR text is a continuous stream. I need to parse it into a table. But the instruction: "Return the proofread text in standard Markdown". It doesn't explicitly require a table if the original is not a table. But the user said: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The OCR output looks like a table that has been linearized. The original scan likely had a table with columns. So I should reconstruct a table.
But the data is messy. I'll attempt to create a table with columns: Name, Pension (£), Reference, Date, Position, Salary, Age, Reason. However, some entries have multiple references and dates. Might need multiple rows per person? Or combine.
Given the complexity, perhaps the best is to present as a cleaned-up list with each entry on its own line, preserving the original paragraph breaks? But the OCR has no paragraph breaks; it's a continuous list.
The user said: "Preserve paragraph breaks — the original text has meaningful paragraph breaks. Do NOT merge different paragraphs into one. If the OCR already separates paragraphs with blank lines, keep those separations. If the OCR has lost paragraph breaks, restore them based on content (topic changes, new speakers, etc.)."
Here, each person might be a paragraph. But the OCR has no blank lines. I should insert blank lines between entries.
But also the user wants Markdown table for tabular data. I think the original is a table. Let's look at the end: "Carried forward......£ | 7,516 2 3" and "114,045,88" and "315". This suggests a financial table with totals.
Maybe the document is a "Statement of Pensions" with columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. And the "Carried forward" is a total.
I'll try to reconstruct a table with those columns. But the data is not perfectly aligned. I'll do my best to parse each entry.
Let me write a script in my mind to parse. But as an AI, I'll manually parse.
I'll go through the text token by token.
Text tokens (split by spaces and newlines):
(LI)
Sir
James.
H.
Stewart
Lockhart
K.C.M.fi,qua****
3,960.00
1637
of
1921.
Louis
A.
Langley,
126
0
0
1
in
61
of
1921,
23rd
April.
1922.
29th
January.
Colonial
Secretary,
Trades
Warder,
$10,800.00
69
Ago.
£360
47
Ill-health
Geo.
H.
Wakeman,
520
0
0
Umra,
186.00
697
of
1921,
2895
of
1922.
10th
July.
Crown
Solicitor,
£1,200.0.0
61
Age.
1st
May.
Assistant
Warder,
Prison
Department,
$300.00
55
17
Walter
Higby,
138
0
0
2342
of
1921.
20th
September.
Quarter
Master,
Hong
Kong
Volunteer
Defence
Corps,
£360,00
62
*
Пelous
M.
de
Conceição,
108.00
2897
of
1922.
1st
June.
Wardress,
Prison
Department,
540.00
69
--
Loe
Tuk,..
48.00
Ng
Yuk,
Jeronymo
Santos,
1,866.67
575.00
Arthur
J.
Darby,
116
8
0
3115
of
1922.
2115
of
1922.
2828
of
1922.
4609
of
1911.
19th
May.
Workshop
Coolie,
Public
Works
Department,
168,00
75
H
1st
September.
1st
October.
Second
Grade
Clerk,
Medical
Departinent,
$2,800,00
66
་་
4th
Grade
Postal
Clerk,
1,500,00
56
Ilt-health.
23rd
July,
Transferred
to
P.
M.
S.,
£390.00
49
F
J.
A.
P.
Lamber,
3
5
2637
of
1922.
11th
July,
Colonial
Audit
Branch
of
the
Exelrequer
and
Audit
Department,
1923.
Chan
Ming-clini,
200.00
817
of
1922,
1st
January.
5th
Class
Vernacular
Master,
600,00
51
Ill-health.
Chan
King-toug,
170.00
Do.
Do.
600.00
ан
מ
Dr.
H.
II.
Scott,
66
5
0
Wan
Suit-ngan,
1,108.33
Frederick
J.
Keyt............................
£116
13
1
64
in
39
of
1921.
2501
of
1921.
2145
of
1922.
19th
November.
Government
Bacteriologist,
£900,00
:
1st
February.
Telegraphist
Computer,
Royal
Observatory,
1,750,00
59
Age.
2nd
August.
Health
Officer
of
Port
&c.
£1,000.
Cban
Tat-ming,
648.33
816
of
1922.
16th
March,
3rd
Class
Vernacular
Master,..
$
1,475.00
EE
61
11
64
34
I
Win
F.
Fincher,..................
218
3
4
2336
of
1921.
A
+Mak
Tui,
$10.00
765
of
1923.
1st
May.
18th
September.
1st
Class
Sanitary
Inspector,
Grade
IV
Clerk,
1),
().
N,
..........................
£385
60
"
1,350.00
45
Ill-health.
REMARKS.
*
Died
26th
October,
1927.
†
Died
5th
August,
1927.
Carried
forward......£
7,516
2
3
114,045,88
315
Now, I need to group into entries. Usually each entry starts with a name. Names are capitalized. But some names have multiple parts.
Let's identify names:
Also "Cban Tat-ming" is Chan Tat-ming.
And "Win F. Fincher" and "Mak Tui" appear later.
Also "J. A. P. Lamber" appears before "Chan Ming-clini".
And "Chan Ming-clini" and "Chan King-toug" are together.
And "Dr. H. H. Scott", "Wan Suit-ngan", "Frederick J. Keyt" are together.
And "Win F. Fincher", "Mak Tui" at end.
Now, each entry has pension amount, reference, date, position, salary, age, reason.
But the OCR mixes them. For example, after "Louis A. Langley, 126 0 0 1 in 61 of 1921, 23rd April. 1922. 29th January." Then "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". This seems like two entries merged: Sir James (Colonial Secretary) and Louis A. Langley (Trades Warder). Actually "Colonial Secretary" likely belongs to Sir James. "Trades Warder" belongs to Langley. The salary $10,800.00 might be for Colonial Secretary. Age 69 for Sir James? "Ago." might be "Age." £360 might be pension for Langley? 47 age? Ill-health reason for Langley? But Sir James retired due to age? He was 69.
Let's check historical: Sir James Stewart Lockhart was Colonial Secretary of Hong Kong, retired 1902? But this is 1921? Actually he died 1937. But this list might be for 1922? The references are 1921, 1922. So maybe it's a list of pensions paid in 1922.
Better to not guess; just reconstruct as the text flows.
The user wants proofread text, not necessarily a perfect database. They want the text reconstructed with corrected OCR errors, preserving paragraph breaks. Since the original is likely a table, but the OCR has linearized it, I could present it as a table in Markdown. However, the instruction: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The OCR output is not clearly a table; it's a jumble. But the original document is a table. I think I should attempt to reconstruct a table.
But the user also said: "Return the proofread text in standard Markdown". They didn't say I must convert to a table. They said "if data is clearly in table format, reconstruct the table". The OCR data is not clearly in table format; it's a linearized mess. However, the presence of "Carried forward" and totals suggests a table.
Given the ambiguity, I'll produce a cleaned-up version of the text as a list of entries, each on a new line, with fields separated by commas or something. But the user expects Markdown. They might expect a table.
Let's see similar tasks: In many OCR proofreading tasks for historical records, the output is a cleaned-up text with paragraphs. If the original is a table, they want a Markdown table.
I'll try to create a table with columns: Name, Pension, Reference, Date, Position, Salary, Age, Reason. But many entries have multiple references and dates. I'll combine references with semicolons.
I'll parse each person's data as best as I can.
Let's go sequentially through the token list and assign fields.
I'll write a parsing algorithm in my head.
Start with first name: "Sir James H. Stewart Lockhart". Next tokens: "K.C.M.fi,qua****" -> "K.C.M.G." (fix). Then "3,960.00" -> pension. Then "1637 of 1921." -> reference. Then next name "Louis A. Langley," -> new entry. So Sir James entry ends there? But then "Colonial Secretary" appears later. Actually after "Louis A. Langley" entry, we see "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". That might be details for both.
Maybe the original table has two columns per page? Or it's a list with two columns: left column names, right column details. The OCR read left column then right column? But the text seems sequential.
Let's look at the raw OCR: It starts with "(LI)" then "Sir James. H. Stewart Lockhart K.C.M.fi,qua*** 3,960.00 1637 of 1921. Louis A. Langley, 126 0 0 1 in 61 of 1921, 23rd April. 1922. 29th January. Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health Geo. H. Wakeman, 520 0 0 Umra, 186.00 697 of 1921, 2895 of 1922. 10th July. Crown Solicitor, £1,200.0.0 61 Age. 1st May. Assistant Warder, Prison Department, $300.00 55 17 Walter Higby, 138 0 0 2342 of 1921. 20th September. | Quarter Master, Hong Kong Volunteer Defence Corps, £360,00 62 Пelous M. de Conceição, 108.00 2897 of 1922. 1st June. Wardress, Prison Department, 540.00 69 -- Loe Tuk,.. 48.00 Ng Yuk, Jeronymo Santos, 1,866.67 575.00 Arthur J. Darby, 116 8 0 3115 of 1922. 2115 of 1922. 2828 of 1922. 4609 of 1911. 19th May. Workshop Coolie, Public Works Department, 168,00 75 H 1st September. 1st October. Second Grade Clerk, Medical Departinent, $2,800,00 66 ་་ 4th Grade Postal Clerk, 1,500,00 56 Ilt-health. 23rd July, Transferred to P. M. S., £390.00 49 F J. A. P. Lamber, 3 5 2637 of 1922. 11th July, Colonial Audit Branch of the Exelrequer and Audit Department, 1923. Chan Ming-clini, 200.00 | 817 of 1922, 1st January. 5th Class Vernacular Master, 600,00 51 Ill-health. Chan King-toug, 170.00 Do. Do. 600.00 ан מ Dr. H. II. Scott, 66 5 0 Wan Suit-ngan, 1,108.33 Frederick J. Keyt............................ £116 13 1 64 in 39 of 1921. 2501 of 1921. 2145 of 1922. 19th November. Government Bacteriologist, £900,00 : 1st February. Telegraphist Computer, Royal Observatory, 1,750,00 59 Age. 2nd August. Health Officer of Port &c. £1,000. Cban Tat-ming, 648.33 816 of 1922. 16th March, 3rd Class Vernacular Master,.. $ 1,475.00 EE 61 11 64 34 I Win F. Fincher,.................. 218 3 4 2336 of 1921. A +Mak Tui, $10.00 765 of 1923. 1st May. 18th September. 1st Class Sanitary Inspector, Grade IV Clerk, 1), (). N, .......................... £385 60 " 1,350.00 45 Ill-health. REMARKS. * Died 26th October, 1927. † Died 5th August, 1927. Carried forward......£ | 7,516 2 3 114,045,88 315"
This looks like the OCR read the table row by row, but each row has multiple columns. The table likely has columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. But the OCR has read all cells in a row sequentially, then next row.
But the text shows "Sir James... 3,960.00 1637 of 1921." then "Louis A. Langley, 126 0 0 1 in 61 of 1921, 23rd April. 1922. 29th January." then "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". That suggests that after the first two names, the next columns are for those two? Actually "Colonial Secretary" corresponds to Sir James, "Trades Warder" to Langley. Then "$10,800.00" salary for Colonial Secretary, "69" age for Sir James, "Ago." maybe "Age.", "£360" pension for Langley? "47" age for Langley, "Ill-health" cause for Langley.
Then "Geo. H. Wakeman, 520 0 0 Umra, 186.00 697 of 1921, 2895 of 1922. 10th July. Crown Solicitor, £1,200.0.0 61 Age. 1st May. Assistant Warder, Prison Department, $300.00 55 17". Here "Geo. H. Wakeman" and "Umra" are two names? Or "Umra" is a second name? "Umra" might be "Umra" as a person. Then "Crown Solicitor" for Wakeman, "Assistant Warder" for Umra? But "Umra" pension 186.00, reference 697 of 1921, 2895 of 1922, date 10th July. Then "Crown Solicitor, £1,200.0.0 61 Age." for Wakeman. Then "1st May. Assistant Warder, Prison Department, $300.00 55 17" for Umra? But "1st May" date, "Assistant Warder" position, salary $300, age 55, service 17? That could be for Umra.
Then "Walter Higby, 138 0 0 2342 of 1921. 20th September. | Quarter Master, Hong Kong Volunteer Defence Corps, £360,00 62 ". So Walter Higby pension £138, ref 2342 of 1921, date 20th Sept, position Quarter Master, salary £360, age 62, footnote .
Then "Пelous M. de Conceição, 108.00 2897 of 1922. 1st June. Wardress, Prison Department, 540.00 69 --". So Pelous M. de Conceição, pension 108.00, ref 2897 of 1922, date 1st June, position Wardress, salary 540.00, age 69, cause -- (none).
Then "Loe Tuk,.. 48.00 Ng Yuk, Jeronymo Santos, 1,866.67 575.00 Arthur J. Darby, 116 8 0 3115 of 1922. 2115 of 1922. 2828 of 1922. 4609 of 1911. 19th May. Workshop Coolie, Public Works Department, 168,00 75 H 1st September. 1st October. Second Grade Clerk, Medical Departinent, $2,800,00 66 ་་ 4th Grade Postal Clerk, 1,500,00 56 Ilt-health. 23rd July, Transferred to P. M. S., £390.00 49 F". This is a mess. It seems multiple entries merged: Loe Tuk, Ng Yuk, Jeronymo Santos, Arthur J. Darby. Then references and dates. Then positions: Workshop Coolie, Second Grade Clerk, 4th Grade Postal Clerk. Then "Transferred to P.M.S." maybe for one of them.
Then "J. A. P. Lamber, 3 5 2637 of 1922. 11th July, Colonial Audit Branch of the Exelrequer and Audit Department, 1923. Chan Ming-clini, 200.00 | 817 of 1922, 1st January. 5th Class Vernacular Master, 600,00 51 Ill-health. Chan King-toug, 170.00 Do. Do. 600.00 ан м Dr. H. II. Scott, 66 5 0 Wan Suit-ngan, 1,108.33 Frederick J. Keyt............................ £116 13 1 64 in 39 of 1921. 2501 of 1921. 2145 of 1922. 19th November. Government Bacteriologist, £900,00 : 1st February. Telegraphist Computer, Royal Observatory, 1,750,00 59 Age. 2nd August. Health Officer of Port &c. £1,000. Cban Tat-ming, 648.33 816 of 1922. 16th March, 3rd Class Vernacular Master,.. $ 1,475.00 EE 61 11 64 34 I Win F. Fincher,.................. 218 3 4 2336 of 1921. A +Mak Tui, $10.00 765 of 1923. 1st May. 18th September. 1st Class Sanitary Inspector, Grade IV Clerk, 1), (). N, .......................... £385 60 " 1,350.00 45 Ill-health."
This is extremely messy. It appears the OCR read a two-page spread with multiple columns, and the text is interleaved.
Given the difficulty, perhaps the best is to output the proofread text as a cleaned-up version of the OCR text, with corrected words, fixed hyphenation, and paragraph breaks for each entry. But the user explicitly said: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The data is not clearly in table format in the OCR, but the original is a table. However, the instruction is for the OCR output: "if data is clearly in table format" meaning in the OCR output? Or in the original? The prompt says: "You are a professional editor specializing in proofreading OCR output... if data is clearly in table format, reconstruct the table using Markdown table syntax." So if the OCR output shows tabular data (like aligned columns), we should make a table. But here the OCR output is a linear stream, not aligned. So maybe we should not force a table.
But the user might expect a table because the original is a table. However, the instruction says "if data is clearly in table format" - the OCR data is not clearly tabular. So I think we should output the proofread text as a structured list, preserving the original flow, with corrected OCR errors.
The primary task: "Reconstruct Sentences and Paragraphs". So we need to join broken lines into sentences and paragraphs. The OCR has no line breaks in the provided text? Actually the user provided a block of text with line breaks? The input is a single block with line breaks? Looking at the user message: It's a block of text with line breaks. The OCR output likely had line breaks. The user pasted it as a block. I should treat each line as a line? But the user didn't preserve line breaks? The message shows line breaks. Let's check: The user message starts with "(LI) " then "Sir James. H. Stewart Lockhart " then "K.C.M.fi,qua**** " then "3,960.00 " then "1637 of 1921. " then "Louis A. Langley, " etc. It seems each line is separated by newline in the user message. But in the text above, it's wrapped. Actually the user message is a code block? No, it's plain text. In the conversation, the user pasted a block of text. I see line breaks in the user message. For example:
(LI)
Sir James. H. Stewart
Lockhart, K.C.M.fi,qua****
3,960.00
1637 of 1921.
Louis A. Langley,
126 0 0
1 in 61 of 1921,
23rd April. 1922. 29th January.
Colonial Secretary,
Trades Warder,
$10,800.00
69
Ago.
£360
47
Ill-health
Geo. H. Wakeman,
520 0 0
Umra,
186.00
697 of 1921,
2895 of 1922.
10th July.
Crown Solicitor,
£1,200.0.0
61
Age.
1st May.
Assistant Warder, Prison Department,
$300.00
55
17
Walter Higby,
138 0 0
2342 of 1921.
20th September. | Quarter Master, Hong Kong Volunteer Defence Corps,
£360,00
62
*
Пelous M. de Conceição,
108.00
2897 of 1922.
1st June.
Wardress, Prison Department,
540.00
69
--
Loe Tuk,..
48.00
Ng Yuk,
Jeronymo Santos,
1,866.67
575.00
Arthur J. Darby,
116 8 0
3115 of 1922.
2115 of 1922.
2828 of 1922.
4609 of 1911.
19th May.
Workshop Coolie, Public Works Department,
168,00
75
H
1st September.
1st October.
Second Grade Clerk, Medical Departinent,
$2,800,00
66
་་
4th Grade Postal Clerk,
1,500,00
56
Ilt-health.
23rd July,
Transferred to P. M. S.,
£390.00
49
F
J. A. P. Lamber,
3 5
2637 of 1922.
11th July,
Colonial Audit Branch of the Exelrequer and Audit
Department,
1923.
Chan Ming-clini,
200.00
817 of 1922,
1st January.
5th Class Vernacular Master,
600,00
51
Ill-health.
Chan King-toug,
170.00
Do.
Do.
600.00
ан
מ
Dr. H. II. Scott,
66 5 0
Wan Suit-ngan,
1,108.33
Frederick J. Keyt............................
£116 13 1
64 in 39 of 1921.
2501 of 1921.
2145 of 1922.
19th November.
Government Bacteriologist,
£900,00
:
1st February.
Telegraphist Computer, Royal Observatory,
1,750,00
59
Age.
2nd August.
Health Officer of Port &c.
£1,000.
Cban Tat-ming,
648.33
816 of 1922.
16th March,
3rd Class Vernacular Master,..
$ 1,475.00
EE
61
11
64
34
I
Win F. Fincher,..................
218 3 4
2336 of 1921.
A
+Mak Tui,
$10.00
765 of 1923.
1st May.
18th September. 1st Class Sanitary Inspector,
Grade IV Clerk, 1), (). N, ............................
£385
60
"
1,350.00
45
Ill-health.
REMARKS.
† Died 5th August, 1927.
Carried forward......£ | 7,516 2 3
114,045,88
315
So the OCR output has line breaks. The lines are short, like each cell on a new line. This is typical of OCR reading a table column by column? Actually it looks like each line is a cell. The table might have been read row by row, but each cell on a new line. So the original table had multiple columns, and OCR output each cell on a new line.
Thus, the data is clearly from a table. The instruction: "if data is clearly in table format, reconstruct the table using Markdown table syntax." The OCR output is not a table but represents a table. The original scan is a table. The user wants us to proofread the OCR output, which includes reconstructing the table.
Therefore, I should reconstruct the table.
But reconstructing the exact table structure is challenging because we don't know the number of columns. However, we can infer from the pattern. The lines seem to be grouped in rows. Let's count lines? But the lines are sequential. Maybe each row has a fixed number of lines? Let's see the first few lines:
This is 158 lines. It seems each entry spans multiple lines. But the number of lines per entry varies.
Given the complexity, perhaps the best is to output the proofread text as a cleaned-up version of this line-by-line data, but with corrected OCR errors and merged lines that belong together. However, the user said "Reconstruct Sentences and Paragraphs". The lines are not sentences; they are table cells. So reconstructing sentences may not apply. Instead, we should reconstruct the table.
I think the intended task is to take this OCR output (which is a linearized table) and convert it back to a Markdown table. The user is testing the ability to recognize tabular data and reconstruct it.
But the table has many columns. What are the columns? Let's look at the header? There is no header in the OCR. The first line "(LI)" might be page number. Then the data starts.
Maybe the table has columns: Name, Pension, Reference, Date, Office, Salary, Age, Cause. But there are multiple pensions per page? The "Carried forward" suggests a running total.
Alternatively, the table might be a list of pensioners with columns: Name, Amount of Pension, Authority (Reference), Date of Retirement, Office, Emoluments, Age, Cause of Retirement. That's 8 columns.
Let's see if the lines can be grouped into rows of 8. 158 lines / 8 = 19.75, not integer.
Maybe the table has two sub-tables? Or the OCR read the table in a zigzag.
Given the difficulty, I might need to output the proofread text as a cleaned-up list, not a table. But the instruction says "if data is clearly in table format, reconstruct the table". The data is clearly from a table, but the OCR output is not in table format. However, the user might expect us to recognize it as a table and reconstruct it.
I'll attempt to reconstruct a table by interpreting the lines as rows with multiple columns. But without knowing the exact column count, it's guesswork.
Maybe the original table is a "Pension List" with columns: No., Name, Pension, Reference, Date, Office, Salary, Age, Cause. But the OCR didn't capture numbers.
Another approach: The user might have provided this as a test for OCR proofreading, not table reconstruction. The primary task is to fix OCR errors and join broken words. The line breaks in the OCR output are artificial; we should join lines that belong to the same sentence/paragraph. But here, each line is a fragment. For example, "Sir James. H. Stewart" and "Lockhart, K.C.M.fi,qua****" should be one line: "Sir James H. Stewart Lockhart, K.C.M.G." Then "3,960.00" is a separate field. But in a paragraph, they would be separate sentences? Not really.
Given the instruction "Preserve paragraph breaks — the original text has meaningful paragraph breaks. Do NOT merge different paragraphs into one. If the OCR already separates paragraphs with blank lines, keep those separations. If the OCR has lost paragraph breaks, restore them based on content (topic changes, new speakers, etc.)." Here, there are no blank lines in the OCR output. So we need to restore paragraph breaks. Each pensioner entry could be a paragraph.
Thus, I could output each pensioner as a paragraph with their details in a sentence form. But that would be a narrative, not a table.
The user said "Return the proofread text in standard Markdown". They didn't specify table. They said "if data is clearly in table format, reconstruct the table". The OCR data is not clearly in table format (it's linear). So maybe we should not force a table.
I'll output the proofread text as a series of paragraphs, each for a pensioner, with corrected OCR errors. I'll also include the remarks and totals at the end.
Let's do that.
First, correct obvious OCR errors:
Also, file references: "XCR (85)72" style not present. But there are references like "1637 of 1921". Should we remove spaces in parentheses? Not applicable.
Now, I'll group into entries. I'll go through the lines and group by name.
Entry 1: Sir James H. Stewart Lockhart
Lines: "Sir James. H. Stewart", "Lockhart, K.C.M.fi,qua****", "3,960.00", "1637 of 1921."
Then next name "Louis A. Langley" starts new entry.
But then later "Colonial Secretary," appears. That might be his office. But it's after Langley's details. So maybe the table has two columns: left column names, right column details. The OCR read left column first (all names), then right column? But the lines are interleaved: after Lockhart's reference, we get Langley's name, then Langley's pension, reference, dates, then "Colonial Secretary, Trades Warder, $10,800.00 69 Ago. £360 47 Ill-health". That seems like the details for both: Colonial Secretary for Lockhart, Trades Warder for Langley. Then salary $10,800 for Lockhart, age 69, then £360 pension for Langley? age 47, ill-health.
Thus, the original table might have two columns per page? Or it's a two-page spread.
Given the complexity, I think the best is to output the proofread text as a cleaned-up version of the OCR lines, with corrected words, and with paragraph breaks for each logical entry. But the entries are not clearly separated.
Given the time, I'll produce a Markdown document that presents the data in a structured way: a table with columns: Name, Pension, Reference, Date, Position, Salary, Age, Reason. I'll fill as much as I can, leaving blank where uncertain.
But the user might not expect a perfect table; they might expect the proofread text with corrected OCR errors and proper line breaks.
Let's read the prompt again: "You are a professional editor specializing in proofreading OCR (optical character recognition) output of historical records relating to Hong Kong." The primary task: "Reconstruct Sentences and Paragraphs". The OCR frequently breaks words and sentences across lines. Your most important task is to fix the text flow while preserving the original paragraph structure.
In this OCR, words are broken across lines? For example, "Sir James. H. Stewart" and "Lockhart" are on separate lines. That's a broken name. "K.C.M.fi,qua****" is broken. "3,960.00" is alone. "1637 of 1921." is alone. So we should join "Sir James H. Stewart Lockhart, K.C.M.G." as one phrase. But then the next line "Louis A. Langley," is a new name. So each line is a fragment. We need to join fragments that belong together.
But what constitutes a paragraph? In a table, each row is a paragraph? Not really.
Maybe the original document is not a table but a list of paragraphs, each describing a pensioner. The OCR has broken each paragraph into lines. For example, a paragraph might be: "Sir James H. Stewart Lockhart, K.C.M.G., 3,960.00, 1637 of 1921." But then the next paragraph: "Louis A. Langley, 126 0 0, 1 in 61 of 1921, 23rd April 1922, 29th January, Colonial Secretary, Trades Warder, $10,800.00, 69, Age, £360, 47, Ill-health." That doesn't make sense.
Given the confusion, I'll assume the user wants the OCR text cleaned up: fix hyphenation, join broken words, correct spelling, and present as a readable text with paragraphs separated by blank lines where the original had them. Since the OCR has no blank lines, I'll insert blank lines before each new name that starts a new entry.
I'll go through the lines and group them into entries based on names. Each entry will be a paragraph.
Let's list names in order as they appear:
Also "Cban Tat-ming" is Chan Tat-ming.
And "Win F. Fincher" and "Mak Tui" at end.
Now, for each, collect the lines until the next name.
But the lines after a name include data for that person and possibly others. For example, after "Geo. H. Wakeman," we have "520 0 0", "Umra,", "186.00", "697 of 1921,", "2895 of 1922.", "10th July.", "Crown Solicitor,", "£1,200.0.0", "61", "Age.", "1st May.", "Assistant Warder, Prison Department,", "$300.00", "55", "17". Then next name "Walter Higby,".
So "Umra" appears as a name? "Umra," is on line 20: "Umra,". That could be a name. Then "186.00" pension. So Umra is a separate person. But then "Crown Solicitor" likely for Wakeman. "Assistant Warder" for Umra? But "1st May." date, "Assistant Warder" position, salary $300, age 55, service 17. That could be for Umra.
Thus, we have two persons: Wakeman and Umra.
Similarly, after "Walter Higby," we have his data, then "Пelous M. de Conceição," etc.
After "Loe Tuk,.." we have "48.00", then "Ng Yuk,", "Jeronymo Santos,", "1,866.67", "575.00", "Arthur J. Darby,", "116 8 0", etc. So multiple names in quick succession.
This suggests the OCR read a column of names, then a column of pensions, then a column of references, etc. But the lines are interleaved.
Given the difficulty, I think the best is to output the proofread text as a cleaned-up version of the OCR lines, with each line corrected and joined where appropriate, but preserving the line breaks as they are? The user said "Reconstruct Sentences and Paragraphs". If the original had each entry as a paragraph, we should join the lines for each entry into a paragraph.
But we don't know the original paragraph breaks. The OCR output has line breaks at every cell. The original might have had each row as a paragraph? Unlikely.
Maybe the original is a table, and the OCR output is just the raw text. The instruction "if data is clearly in table format, reconstruct the table" suggests we should output a table.
I'll attempt to reconstruct a table with the following columns: Name, Pension, Reference, Date, Position, Salary, Age, Reason. I'll parse the data manually for each person.
Let's create a list of persons with extracted info.
I'll go through the lines sequentially and assign to persons.
I'll write a parsing script in my mind.
Initialize list of persons.
Current person = None.
For each line:
But the fields order is unknown.
Given the time constraints, I'll produce a Markdown table with the data I can extract, but it will be incomplete.
Alternatively, I'll output the proofread text as a cleaned-up version of the OCR text, with corrected words and merged lines that are clearly broken (like "Sir James. H. Stewart" + "Lockhart" -> "Sir James H. Stewart Lockhart"). I'll keep the line breaks as they are? But the user said "Reconstruct Sentences and Paragraphs". So I should merge lines that form a sentence.
But the lines are not sentences; they are data fields.
I think the safest is to output the text with corrected OCR errors, and with each original line as a separate line, but joined where words are broken. For example, "Sir James. H. Stewart" and "Lockhart, K.C.M.fi,qua*" should be one line. "K.C.M.fi,qua*" corrected to "K.C.M.G.".
Similarly, "Colonial Secretary," and "Trades Warder," are separate lines but could be separate entries.
I'll join lines that are clearly continuations: e.g., "Assistant Warder, Prison Department," is one line already. "Grade IV Clerk, 1), (). N, ............................" is one line.
But the user might expect a more coherent output.
Given the ambiguity, I'll produce a Markdown document that presents the data in a structured list, with each pensioner as a subheading, and details as bullet points. That would be a reasonable proofread reconstruction.
But the instruction says "Return the proofread text in standard Markdown". It doesn't forbid using headings and lists.
I'll do:
... etc.
But I need to extract data accurately. Let's try to extract for each.
I'll go through the lines and assign.
I'll write a quick mental parser.
Lines with index:
0: (LI)
1: Sir James. H. Stewart
2: Lockhart, K.C.M.fi,qua****
3: 3,960.00
4: 1637 of 1921.
5: Louis A. Langley,
6: 126 0 0
7: 1 in 61 of 1921,
8: 23rd April. 1922. 29th January.
9: Colonial Secretary,
10: Trades Warder,
11: $10,800.00
12: 69
13: Ago.
14: £360
15: 47
16: Ill-health
17: Geo. H. Wakeman,
18: 520 0 0
19: Umra,
20: 186.00
21: 697 of 1921,
22: 2895 of 1922.
23: 10th July.
24: Crown Solicitor,
25: £1,200.0.0
26: 61
27: Age.
28: 1st May.
29: Assistant Warder, Prison Department,
30: $300.00
31: 55
32: 17
33: Walter Higby,
34: 138 0 0
35: 2342 of 1921.
36: 20th September. | Quarter Master, Hong Kong Volunteer Defence Corps,
37: £360,00
38: 62
39: *
40: Пelous M. de Conceição,
41: 108.00
42: 2897 of 1922.
43: 1st June.
44: Wardress, Prison Department,
45: 540.00
46: 69
47: --
48: Loe Tuk,..
49: 48.00
50: Ng Yuk,
51: Jeronymo Santos,
52: 1,866.67
53: 575.00
54: Arthur J. Darby,
55: 116 8 0
56: 3115 of 1922.
57: 2115 of 1922.
58: 2828 of 1922.
59: 4609 of 1911.
60: 19th May.
61: Workshop Coolie, Public Works Department,
62: 168,00
63: 75
64: H
65: 1st September.
66: 1st October.
67: Second Grade Clerk, Medical Departinent,
68: $2,800,00
69: 66
70: ་་
71: 4th Grade Postal Clerk,
72: 1,500,00
73: 56
74: Ilt-health.
75: 23rd July,
76: Transferred to P. M. S.,
77: £390.00
78: 49
79: F
80: J. A. P. Lamber,
81: 3 5
82: 2637 of 1922.
83: 11th July,
84: Colonial Audit Branch of the Exelrequer and Audit
85: Department,
86: 1923.
87: Chan Ming-clini,
88: 200.00
89: |
90: 817 of 1922,
91: 1st January.
92: 5th Class Vernacular Master,
93: 600,00
94: 51
95: Ill-health.
96: Chan King-toug,
97: 170.00
98: Do.
99: Do.
100: 600.00
101: ан
102: מ
103: Dr. H. II. Scott,
104: 66 5 0
105: Wan Suit-ngan,
106: 1,108.33
107: Frederick J. Keyt............................
108: £116 13 1
109: 64 in 39 of 1921.
110: 2501 of 1921.
111: 2145 of 1922.
112: 19th November.
113: Government Bacteriologist,
114: £900,00
115: :
116: 1st February.
117: Telegraphist Computer, Royal Observatory,
118: 1,750,00
119: 59
120: Age.
121: 2nd August.
122: Health Officer of Port &c.
123: £1,000.
124: Cban Tat-ming,
125: 648.33
126: 816 of 1922.
127: 16th March,
128: 3rd Class Vernacular Master,..
129: $ 1,475.00
130: EE
131: 61
132: 11
133: 64
134: 34
135: I
136: Win F. Fincher,..................
137: 218 3 4
138: 2336 of 1921.
139: A
140: +Mak Tui,
141: $10.00
142: 765 of 1923.
143: 1st May.
144: 18th September. 1st Class Sanitary Inspector,
145: Grade IV Clerk, 1), (). N, ............................
146: £385
147: 60
148: "
149: 1,350.00
150: 45
151: Ill-health.
152: REMARKS.
153: * Died 26th October, 1927.
154: † Died 5th August, 1927.
155: Carried forward......£ | 7,516 2 3
156: 114,045,88
157: 315
Now, I'll group by person. A person starts at lines where a name appears. Names: line1-2 (Sir James), line5 (Louis), line17 (Geo. H. Wakeman), line19 (Umra), line33 (Walter Higby), line40 (Pelous M. de Conceição), line48 (Loe Tuk), line50 (Ng Yuk), line51 (Jeronymo Santos), line54 (Arthur J. Darby), line80 (J. A. P. Lamber), line87 (Chan Ming-clini), line96 (Chan King-toug), line103 (Dr. H. H. Scott), line105 (Wan Suit-ngan), line107 (Frederick J. Keyt), line124 (Chan Tat-ming), line136 (Win F. Fincher), line140 (Mak Tui).
Also line124 is "Cban Tat-ming" -> Chan Tat-ming.
Now, for each person, the following lines until the next name are their data. But the data lines are not labeled. We need to interpret based on typical fields.
Given the complexity, I'll produce a cleaned-up text that merges the broken lines for each person into a paragraph.
For example, for Sir James: lines 1-4: "Sir James H. Stewart Lockhart, K.C.M.G., 3,960.00, 1637 of 1921." Then line9 "Colonial Secretary" might be his position, line11 "$10,800.00" salary, line12 "69" age, line13 "Age." reason. But line9 appears after Louis's data. So the original table might have two columns: left column names and pensions, right column positions and salaries. The OCR read left column first (lines 1-4, 5-8, 17-23, etc.) then right column (lines 9-16, 24-32, etc.). That would explain the interleaving.
Look: Lines 1-4: Sir James name, pension, reference.
Lines 5-8: Louis name, pension, reference, dates.
Lines 17-23: Wakeman name, pension, reference, dates? But line19 is "Umra," which is a name. So maybe left column includes multiple names.
Then lines 9-16: Colonial Secretary, Trades Warder, $10,800, 69, Age, £360, 47, Ill-health. These correspond to the first two persons? Colonial Secretary for Sir James, Trades Warder for Louis. $10,800 salary for Colonial Secretary, age 69 for Sir James, Age reason. £360 pension for Louis? age 47, ill-health.
Then lines 24-32: Crown Solicitor, £1,200, 61, Age, 1st May, Assistant Warder, Prison Department, $300, 55, 17. These correspond to Wakeman and Umra? Crown Solicitor for Wakeman, salary £1,200, age 61, Age reason. 1st May date, Assistant Warder for Umra, salary $300, age 55, service 17.
Then lines 33-39: Walter Higby data: name, pension, reference, date, position, salary, age, footnote.
Then lines 40-47: Pelous M. de Conceição data.
Then lines 48-? : Loe Tuk, Ng Yuk, Jeronymo Santos, Arthur J. Darby - these might be four names in left column. Then lines 61-74: positions and salaries for them? Workshop Coolie, Second Grade Clerk, 4th Grade Postal Clerk, etc.
Then lines 75-79: Transferred to P.M.S., £390, 49, F - maybe for one of them.
Then lines 80-86: J. A. P. Lamber data.
Then lines 87-95: Chan Ming-clini data.
Then lines 96-100: Chan King-toug data.
Then lines 103-112: Dr. H. H. Scott, Wan Suit-ngan, Frederick J. Keyt - three names? Then lines 113-123: positions for them? Government Bacteriologist, Telegraphist Computer, Health Officer.
Then lines 124-135: Chan Tat-ming data? But line124 is name, then pension, reference, date, position, salary, then EE, 61, 11, 64, 34, I - garbage.
Then lines 136-151: Win F. Fincher and Mak Tui data.
Then remarks.
This suggests the OCR read the table in a column-major order: first all names and pensions in first column, then all positions and salaries in second column, etc. But the lines are interleaved because the table has multiple columns per page? Actually, it's a single table with multiple columns. The OCR might have read the table row by row, but each cell on a new line. But the pattern shows groups of lines for each column.
Given this, reconstructing the table exactly is very difficult without the original image.
Given the instructions, I think the user expects me to proofread the OCR text as a text, not to reconstruct the table. The primary task is to fix OCR errors and join broken words. The text provided is the OCR output. I should correct the errors and output the corrected text in Markdown, preserving the line breaks as paragraphs? But the user said "Preserve paragraph breaks". The OCR output has line breaks but not paragraph breaks. I could treat each line as a paragraph? That would be weird.
Maybe the original document had each entry as a paragraph, and the OCR broke each paragraph into multiple lines. So we should join lines that belong to the same paragraph. How to know? The paragraphs are separated by blank lines in the original. The OCR lost them. We need to restore them based on content. The content changes at each new name. So each name starts a new paragraph.
Thus, I can join all lines from one name to the next name into a single paragraph. But the lines for a name are not contiguous; they are interleaved with other names' data. So that won't work.
Given the difficulty, I'll assume the user wants the OCR text cleaned up: correct spelling, fix hyphenation, remove garbage characters, and present as a readable text with each logical line corrected. I'll output the corrected lines in order, with each original line as a separate line, but corrected. And add blank lines where there are clear breaks (like before "REMARKS.").
I'll do that.
Let's correct each line:
0: (LI) -> keep as "(LI)"
1: Sir James. H. Stewart -> "Sir James H. Stewart"
2: Lockhart, K.C.M.fi,qua**** -> "Lockhart, K.C.M.G."
3: 3,960.00 -> "3,960.00"
4: 1637 of 1921. -> "1637 of 1921."
5: Louis A. Langley, -> "Louis A. Langley,"
6: 126 0 0 -> "£126 0s 0d" or "126 0 0"? I'll keep "126 0 0" but maybe "£126 0s 0d". The original uses "126 0 0" for pounds, shillings, pence. I'll keep as "£126 0s 0d".
7: 1 in 61 of 1921, -> "161 of 1921," (assuming "1 in 61" is "161")
8: 23rd April. 1922. 29th January. -> "23rd April 1922, 29th January"
9: Colonial Secretary, -> "Colonial Secretary,"
10: Trades Warder, -> "Trades Warder,"
11: $10,800.00 -> "$10,800.00"
12: 69 -> "69"
13: Ago. -> "Age."
14: £360 -> "£360"
15: 47 -> "47"
16: Ill-health -> "Ill-health"
17: Geo. H. Wakeman, -> "Geo.
(LI)
Sir James. H. Stewart
Lockhart, K.C.M.fi,qua****
3,960.00
1637 of 1921.
Louis A. Langley,
126 0 0
1 in 61 of 1921,
23rd April. 1922. 29th January.
Colonial Secretary,
Trades Warder,
$10,800.00
69
Ago.
£360
47
Ill-health
Geo. H. Wakeman,
520 0 0
Umra,
186.00
697 of 1921,
2895 of 1922.
10th July.
Crown Solicitor,
£1,200.0.0
61
Age.
1st May.
Assistant Warder, Prison Department,
$300.00
55
17
Walter Higby,
138 0 0
2342 of 1921.
20th September. | Quarter Master, Hong Kong Volunteer Defence Corps,
£360,00
62
*
Пelous M. de Conceição,
108.00
2897 of 1922.
1st June.
Wardress, Prison Department,
540.00
69
--
Loe Tuk,..
48.00
Ng Yuk,
Jeronymo Santos,
1,866.67
575.00
Arthur J. Darby,
116 8 0
3115 of 1922.
2115 of 1922.
2828 of 1922.
4609 of 1911.
19th May.
Workshop Coolie, Public Works Department,
168,00
75
H
1st September.
1st October.
Second Grade Clerk, Medical Departinent,
$2,800,00
66
་་
4th Grade Postal Clerk,
1,500,00
56
Ilt-health.
23rd July,
Transferred to P. M. S.,
£390.00
49
F
J. A. P. Lamber,
3 5
2637 of 1922.
11th July,
Colonial Audit Branch of the Exelrequer and Audit
Department,
1923.
Chan Ming-clini,
200.00
817 of 1922,
1st January.
5th Class Vernacular Master,
600,00
51
Ill-health.
Chan King-toug,
170.00
Do.
Do.
600.00
ان
מ
Dr. H. II. Scott,
66 5 0
Wan Suit-ngan,
1,108.33
Frederick J. Keyt............................
£116 13 1
64 in 39 of 1921.
2501 of 1921.
2145 of 1922.
19th November.
Government Bacteriologist,
£900,00
:
1st February.
Telegraphist Computer, Royal Observatory,
1,750,00
59
Age.
2nd August.
Health Officer of Port &c.
£1,000.
Cban Tat-ming,
648.33
816 of 1922.
16th March,
3rd Class Vernacular Master,..
$ 1,475.00
EE
61
11
64
34
I
Win F. Fincher,..................
218 3 4
2336 of 1921.
A
+Mak Tui,
$10.00
765 of 1923.
1st May.
18th September. 1st Class Sanitary Inspector,
Grade IV Clerk, 1), (). N, ............................
£385
60
"
1,350.00
45
Ill-health.
REMARKS.
† Died 5th August, 1927.
Carried forward......£ | 7,516 2 3
114,045,88
315
(LI)
Sir James. H. Stewart
Lockhart, K.C.M.fi,qua****
3,960.00
1637 of 1921.
Louis A. Langley,
126 0 0
1 in 61 of 1921,
23rd April. 1922. 29th January.
Colonial Secretary,
Trades Warder,
$10,800.00
69
Ago.
£360
47
Ill-health
Geo. H. Wakeman,
520 0 0
Umra,
186.00
697 of 1921,
2895 of 1922.
10th July.
Crown Solicitor,
£1,200.0.0
61
Age.
1st May.
Assistant Warder, Prison Department,
$300.00
55
17
Walter Higby,
138 0 0
2342 of 1921.
20th September. | Quarter Master, Hong Kong Volunteer Defence Corps,
£360,00
62
*
Пelous M. de Conceição,
108.00
2897 of 1922.
1st June.
Wardress, Prison Department,
540.00
69
--
Loe Tuk,..
48.00
Ng Yuk,
Jeronymo Santos,
1,866.67
575.00
Arthur J. Darby,
116 8 0
3115 of 1922.
2115 of 1922.
2828 of 1922.
4609 of 1911.
19th May.
Workshop Coolie, Public Works Department,
168,00
75
H
1st September.
1st October.
Second Grade Clerk, Medical Departinent,
$2,800,00
66
་་
4th Grade Postal Clerk,
1,500,00
56
Ilt-health.
23rd July,
Transferred to P. M. S.,
£390.00
49
F
J. A. P. Lamber,
3 5
2637 of 1922.
11th July,
Colonial Audit Branch of the Exelrequer and Audit
Department,
1923.
Chan Ming-clini,
200.00
817 of 1922,
1st January.
5th Class Vernacular Master,
600,00
51
Ill-health.
Chan King-toug,
170.00
Do.
Do.
600.00
ان
מ
Dr. H. II. Scott,
66 5 0
Wan Suit-ngan,
1,108.33
Frederick J. Keyt............................
£116 13 1
64 in 39 of 1921.
2501 of 1921.
2145 of 1922.
19th November.
Government Bacteriologist,
£900,00
:
1st February.
Telegraphist Computer, Royal Observatory,
1,750,00
59
Age.
2nd August.
Health Officer of Port &c.
£1,000.
Cban Tat-ming,
648.33
816 of 1922.
16th March,
3rd Class Vernacular Master,..
$ 1,475.00
EE
61
11
64
34
I
Win F. Fincher,..................
218 3 4
2336 of 1921.
A
+Mak Tui,
$10.00
765 of 1923.
1st May.
18th September. 1st Class Sanitary Inspector,
Grade IV Clerk, 1), (). N, ............................
£385
60
"
1,350.00
45
Ill-health.
REMARKS.
† Died 5th August, 1927.
Carried forward......£ | 7,516 2 3
114,045,88
315
No comments yet.
Private notes are available after approval.