The user wants me to proofread OCR output of a historical Hong Kong document. The text appears to be a comparative yearly statement of tax/duty collections for 1912 and 1913. I need to reconstruct the text in Markdown format, preserving paragraph structure, fixing OCR errors, and reconstructing tables.
Looking at the OCR text, it seems to be a table with columns: "Specify each separate Tax or Duty according to the classification employed in the Estimates", "Amount collected in the Year 1912", "Amount collected in the Year 1913". The OCR has broken the table structure.
I need to reconstruct this as a proper Markdown table. The text shows various license types and fees with amounts for 1912 and 1913. There are also section headers like "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued" and "1. FEES OF COURT OR OFFICE, PAYMENTS FOR SPECIFIC PURPOSES, AND REIMBURSEMENTS IN AID :-".
I should create a table with three columns: Description, 1912 Amount, 1913 Amount. The OCR has some formatting issues: numbers like "3,114,758,99" should be "3,114,758.99" (decimal point). Also "3,338,304,97" → "3,338,304.97". Some amounts have extra spaces or missing decimal points.
I'll parse the data line by line. The OCR seems to have mixed up columns. Let me try to reconstruct:
First line: "Brought forward.$ 3,114,758,99 3,338,304,97" → Brought forward, $3,114,758.99, $3,338,304.97
Then "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued" is a section header.
Then items:
Let me read the OCR more carefully. It seems the OCR read columns incorrectly. The original likely had two columns of amounts (1912 and 1913) for each item. The OCR output shows numbers in a single column but they alternate? Actually looking at the text:
"Money Changers' Liceuces,
Opium Fario,
Pawnbrokers' Licences,..
Permits to cut Earth and Stone, New Territories,.
Pharmacy Ordinance Licences, - • • •
Pilots' Licences,.
Prospecting Licences, N.T.,.......................
Sandalwood Mill Licences, N.T.,
Special Fruit Licences,....
Stamp Duties,.........
Steam-Launch Licences,
Theatrical Licences,
Water-Wheel Licences,....
PURPOSES, AND REIMBURSEMENTS IN AID :-
3,338,304,97
4,110.25
1,183,200.00
3,260.00
1,183,200.00
92,362.50
104,240.82
258.00
268.00
20.00
20.00
70.00
70.00
500.00
2,000.00
10,00
4,222.53
1,151,311.88
4,769.55
1,065,914.90
.4,950.00
5,051.50
1,515.00
1,715.00
34,00
44.00
Analyses,
10,260.50
8,816.50
Bills of Health,
Births and Deaths Registration,
Cemetery Fees,....
Certificates to Chiuese eatering America,
Chemists and Druggists Registration...................
Chinese Cemetery Fees,
Companies Registration,....................
Contribution from Chinese Dispensaries, &c., .............................................................
8,358.50
10,044.00
1,711,80
1,987.10
931.00
655,98
425,00
225.00
100.00
150,00
2,611.50
2,794.00
9,301.30
9,276.60
389,94
210.00
·Coutribution from Wei-hai-wei Government for Clerical Assistance,
Au
360.00
430.00
Court Fees,
39,679.25
41,847.77
Crown Lensus,
1.820.00
3,670.00
Deeds Registration,
43,408.95
51,437.50
Engagement and Discharge of Seamen,
23,784.20
81,576.20
Carried forward,...........................................
5,700,404.19
5,872,039.39"
This is very messy. It appears the OCR has lost the table structure and just output text sequentially. I need to infer the correct pairing of items with amounts.
Given the instruction to reconstruct tables using Markdown table syntax, I should create a table with columns: "Tax/Duty", "1912 Amount", "1913 Amount". I'll need to match each description with its two amounts.
Let me try to parse logically. The first section: "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued" with items listed. Then amounts follow. But the amounts seem to be in two columns interleaved? Actually, the OCR might have read the table row by row but the columns got mixed.
Better approach: The original document is a "COMPARATIVE YEARLY STATEMENT" with columns: "Specify each separate Tax or Duty...", "Amount collected in the Year 1912", "Amount collected in the Year 1913". So each row has three cells.
The OCR output shows descriptions and then numbers. Perhaps the numbers are in order: first all 1912 amounts, then all 1913 amounts? But that seems unlikely.
Let me look at the numbers: After "Water-Wheel Licences,...." there is a blank line then "1. FEES OF COURT OR OFFICE..." then "3,338,304,97" (which is likely the brought forward total for 1913? Actually earlier "Brought forward.$ 3,114,758,99 3,338,304,97" - that's two numbers: 1912 and 1913 brought forward totals.
Then after the section header, there is a list of numbers: 4,110.25; 1,183,200.00; 3,260.00; 1,183,200.00; 92,362.50; 104,240.82; 258.00; 268.00; 20.00; 20.00; 70.00; 70.00; 500.00; 2,000.00; 10,00; 4,222.53; 1,151,311.88; 4,769.55; 1,065,914.90; .4,950.00; 5,051.50; 1,515.00; 1,715.00; 34,00; 44.00; then "Analyses," then 10,260.50; 8,816.50; "Bills of Health,"; "Births and Deaths Registration,"; "Cemetery Fees,...."; "Certificates to Chiuese eatering America,"; "Chemists and Druggists Registration..................."; "Chinese Cemetery Fees,"; "Companies Registration,...................."; "Contribution from Chinese Dispensaries, &c., ............................................................."; then numbers: 8,358.50; 10,044.00; 1,711,80; 1,987.10; 931.00; 655,98; 425,00; 225.00; 100.00; 150,00; 2,611.50; 2,794.00; 9,301.30; 9,276.60; 389,94; 210.00; "·Coutribution from Wei-hai-wei Government for Clerical Assistance,"; "Au"; 360.00; 430.00; "Court Fees,"; 39,679.25; 41,847.77; "Crown Lensus,"; 1.820.00; 3,670.00; "Deeds Registration,"; 43,408.95; 51,437.50; "Engagement and Discharge of Seamen,"; 23,784.20; 81,576.20; "Carried forward,..........................................."; 5,700,404.19; 5,872,039.39.
This is extremely messy. It seems the OCR has interleaved descriptions and numbers incorrectly. Some descriptions appear after numbers. For example, "Analyses," appears between numbers 44.00 and 10,260.50. Then "Bills of Health," appears after 8,816.50. Then "Births and Deaths Registration," after that. Then "Cemetery Fees,...." then "Certificates to Chiuese eatering America," then "Chemists and Druggists Registration..................." then "Chinese Cemetery Fees," then "Companies Registration,...................." then "Contribution from Chinese Dispensaries, &c., ............................................................." then numbers.
It appears that the OCR read the table column by column? Or maybe the original had two pages side by side? The user said: "Restore column reading order — if text was originally in multiple columns but OCR read it in the wrong order, reorder it to make logical sense (left-to-right, top-to-bottom)."
Given the complexity, I might need to make a best-effort reconstruction. However, the instruction says: "Do not add or remove substantive words. Do not rephrase, rewrite, or change the author's style, tone, or word choice. Only correct what is clearly wrong." But also "Reconstruct tabular data using Markdown table syntax."
I think the best is to present the data as a table with the descriptions and amounts as they appear, but I need to pair them correctly. Since I cannot know the exact original pairing, I might need to infer from typical Hong Kong revenue categories.
But perhaps the OCR output is actually a single column of text that includes both the descriptions and the amounts in order: description, 1912 amount, 1913 amount? Let's test: "Money Changers' Liceuces," then next line "Opium Fario," then "Pawnbrokers' Licences,.." then "Permits to cut Earth and Stone, New Territories,." then "Pharmacy Ordinance Licences, - • • • " then "Pilots' Licences,." then "Prospecting Licences, N.T.,......................." then "Sandalwood Mill Licences, N.T.," then "Special Fruit Licences,...." then "Stamp Duties,........." then "Steam-Launch Licences," then "Theatrical Licences," then "Water-Wheel Licences,...." then "1. FEES OF COURT OR OFFICE, PAYMENTS FOR SPECIFIC PURPOSES, AND REIMBURSEMENTS IN AID :-" then "3,338,304,97" (maybe this is the 1913 brought forward? but already had brought forward). Then "4,110.25" (maybe Money Changers' 1912), "1,183,200.00" (Opium 1912?), "3,260.00" (Pawnbrokers 1912?), "1,183,200.00" (Permits 1912?), "92,362.50" (Pharmacy 1912?), "104,240.82" (Pilots 1912?), "258.00" (Prospecting 1912?), "268.00" (Sandalwood 1912?), "20.00" (Special Fruit 1912?), "20.00" (Stamp Duties 1912?), "70.00" (Steam-Launch 1912?), "70.00" (Theatrical 1912?), "500.00" (Water-Wheel 1912?), "2,000.00" (maybe something else), "10,00" (maybe 10.00), "4,222.53" (Money Changers 1913?), "1,151,311.88" (Opium 1913?), "4,769.55" (Pawnbrokers 1913?), "1,065,914.90" (Permits 1913?), ".4,950.00" (Pharmacy 1913?), "5,051.50" (Pilots 1913?), "1,515.00" (Prospecting 1913?), "1,715.00" (Sandalwood 1913?), "34,00" (Special Fruit 1913?), "44.00" (Stamp Duties 1913?), then "Analyses," (new item), "10,260.50" (Analyses 1912?), "8,816.50" (Analyses 1913?), "Bills of Health," (new item), "Births and Deaths Registration," (new item), "Cemetery Fees,...." (new item), "Certificates to Chiuese eatering America," (new item), "Chemists and Druggists Registration..................." (new item), "Chinese Cemetery Fees," (new item), "Companies Registration,...................." (new item), "Contribution from Chinese Dispensaries, &c., ............................................................." (new item), then numbers: "8,358.50" (maybe Analyses 1912? but already had), "10,044.00" (maybe Bills of Health 1912?), "1,711,80" (Births 1912?), "1,987.10" (Cemetery 1912?), "931.00" (Certificates 1912?), "655,98" (Chemists 1912?), "425,00" (Chinese Cemetery 1912?), "225.00" (Companies 1912?), "100.00" (Contribution 1912?), "150,00" (maybe next), "2,611.50" (Wei-hai-wei 1912?), "2,794.00" (Court Fees 1912?), "9,301.30" (Crown Leases 1912?), "9,276.60" (Deeds 1912?), "389,94" (Engagement 1912?), "210.00" (Carried forward 1912?), then "·Coutribution from Wei-hai-wei Government for Clerical Assistance," (description), "Au" (maybe typo), "360.00" (1912?), "430.00" (1913?), "Court Fees," (description), "39,679.25" (1912?), "41,847.77" (1913?), "Crown Lensus," (description), "1.820.00" (1912?), "3,670.00" (1913?), "Deeds Registration," (description), "43,408.95" (1912?), "51,437.50" (1913?), "Engagement and Discharge of Seamen," (description), "23,784.20" (1912?), "81,576.20" (1913?), "Carried forward,..........................................." (description), "5,700,404.19" (1912?), "5,872,039.39" (1913?).
This is too speculative. The OCR is too garbled. However, the user expects me to proofread and reconstruct. Perhaps the original image is a table with two columns of amounts. The OCR has read it as a single stream. I need to use my best judgment to create a plausible table.
Given the instruction: "If data is clearly in table format, reconstruct the table using Markdown table syntax." I should create a table with the headers: "Tax/Duty", "1912 Amount", "1913 Amount". Then list each item with its two amounts as they appear in the OCR, but I need to match them.
Maybe the OCR output is actually the raw text from a two-column layout where the left column is descriptions and the right column is amounts for 1912 and 1913 side by side? But the OCR has mixed them.
Another approach: The text starts with "(08)" then "Specify each separate Tax or Duty according to the classification employed in the Estimates," then "COMPARATIVE YEARLY STATE-" then "Amount collected Amount collected in the Year 1912. |in the Year 1913." then "G." then "Brought forward.$ 3,114,758,99 3,338,304,97". So the first row is "Brought forward" with amounts 3,114,758.99 and 3,338,304.97.
Then "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued" is a section header.
Then a list of items: Money Changers' Licences, Opium Farm, Pawnbrokers' Licences, Permits to cut Earth and Stone, New Territories, Pharmacy Ordinance Licences, Pilots' Licences, Prospecting Licences, N.T., Sandalwood Mill Licences, N.T., Special Fruit Licences, Stamp Duties, Steam-Launch Licences, Theatrical Licences, Water-Wheel Licences.
Then "1. FEES OF COURT OR OFFICE, PAYMENTS FOR SPECIFIC PURPOSES, AND REIMBURSEMENTS IN AID :-" another section header.
Then a list of items: Analyses, Bills of Health, Births and Deaths Registration, Cemetery Fees, Certificates to Chinese entering America, Chemists and Druggists Registration, Chinese Cemetery Fees, Companies Registration, Contribution from Chinese Dispensaries, &c., Contribution from Wei-hai-wei Government for Clerical Assistance, Court Fees, Crown Leases, Deeds Registration, Engagement and Discharge of Seamen.
Then "Carried forward" with totals.
The amounts are interspersed. Perhaps the amounts for the first section are the numbers before "Analyses,"? Let's count items in first section: 13 items (Money Changers to Water-Wheel). The numbers before "Analyses," are: 4,110.25; 1,183,200.00; 3,260.00; 1,183,200.00; 92,362.50; 104,240.82; 258.00; 268.00; 20.00; 20.00; 70.00; 70.00; 500.00; 2,000.00; 10,00; 4,222.53; 1,151,311.88; 4,769.55; 1,065,914.90; .4,950.00; 5,051.50; 1,515.00; 1,715.00; 34,00; 44.00. That's 26 numbers. 13 items * 2 = 26. Perfect! So the first 13 numbers are 1912 amounts, the next 13 are 1913 amounts. Let's verify: 13 items, 26 numbers. The first 13: 4,110.25; 1,183,200.00; 3,260.00; 1,183,200.00; 92,362.50; 104,240.82; 258.00; 268.00; 20.00; 20.00; 70.00; 70.00; 500.00. The next 13: 2,000.00; 10,00; 4,222.53; 1,151,311.88; 4,769.55; 1,065,914.90; .4,950.00; 5,051.50; 1,515.00; 1,715.00; 34,00; 44.00. Wait that's 12? Let's count: after 500.00, the next numbers: 2,000.00 (1), 10,00 (2), 4,222.53 (3), 1,151,311.88 (4), 4,769.55 (5), 1,065,914.90 (6), .4,950.00 (7), 5,051.50 (8), 1,515.00 (9), 1,715.00 (10), 34,00 (11), 44.00 (12). That's only 12. But we need 13. Maybe the "2,000.00" belongs to the previous? Or maybe there is an extra item. Let's recount items:
That's 13. The numbers before "Analyses," are 26 numbers? Let's list all numbers from "3,338,304,97" to "44.00" sequentially:
Then after section header, the numbers:
That's 25 numbers after the brought forward? Actually 25 numbers from 4,110.25 to 44.00. But we need 26 for 13 items. Maybe the "2,000.00" and "10,00" are part of the same item? Or maybe there is an extra item like "Opium Farm" might have two sub-items? Hmm.
Alternatively, the first 13 numbers (2-14) are 1912 amounts for the 13 items. The next 13 numbers (15-27) would be 1913 amounts, but we only have 12 numbers (15-26). Number 27 might be missing or the "Analyses," appears before the last number. But "Analyses," is a description for the next section. So maybe the 13th 1913 amount is missing due to OCR. Or maybe the "2,000.00" is not an amount but something else.
Let's check the amounts: For Money Changers' Licences, typical amounts might be around 4,000. So 4,110.25 (1912) and 4,222.53 (1913) plausible. Opium Farm: huge revenue, 1,183,200.00 (1912) and 1,151,311.88 (1913) plausible. Pawnbrokers: 3,260.00 and 4,769.55 plausible. Permits: 1,183,200.00 and 1,065,914.90 plausible. Pharmacy: 92,362.50 and .4,950.00? That's 4,950.00? But 92k vs 4.9k seems off. Maybe Pharmacy is 92,362.50 and 5,051.50? Pilots: 104,240.82 and 1,515.00? That's a big drop. Prospecting: 258.00 and 1,715.00. Sandalwood: 268.00 and 34.00? Special Fruit: 20.00 and 44.00? Stamp Duties: 20.00 and ?. Steam-Launch: 70.00 and ?. Theatrical: 70.00 and ?. Water-Wheel: 500.00 and ?. Then there are extra numbers 2,000.00 and 10.00. Maybe those are for other items not listed? Or maybe the list includes "Opium Farm" and "Opium Licences" separate? The OCR says "Opium Fario" maybe "Opium Farm". Could be two opium items.
Given the difficulty, I think the best is to present the table as reconstructed from the OCR with the data aligned as best as possible, noting uncertainties. But the instruction says "Do not add or remove substantive words." and "Only correct what is clearly wrong." I should not guess missing data. However, I must reconstruct the table. Perhaps I can create a table with the descriptions and the amounts as they appear in the OCR in two columns, but I need to pair them correctly.
Given the time, I'll produce a Markdown table with the sections and items, using the amounts that seem to correspond. I'll use the first 13 numbers as 1912 and the next 13 as 1913, but since there are only 12 in the second group, I'll assume the last one is missing and put "..." for missing.
But the user might expect a clean reconstruction. Let me think: The OCR text is from a known historical document: "Hong Kong Blue Book" or "Estimates". The comparative yearly statement of revenue. I could search my memory but I don't have exact data.
Given the constraints, I'll output the text as a series of paragraphs and a table for the tabular data. I'll keep the section headers as markdown headers. For the table, I'll create a markdown table with columns: "Tax/Duty", "1912", "1913". I'll populate with the data I can parse.
I'll start with the first paragraph: "(08) Specify each separate Tax or Duty according to the classification employed in the Estimates, COMPARATIVE YEARLY STATEMENT Amount collected in the Year 1912. | Amount collected in the Year 1913. G. Brought forward. $3,114,758.99 $3,338,304.97"
Then section "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued" as a subheader.
Then a table for those licences.
Then section "1. FEES OF COURT OR OFFICE, PAYMENTS FOR SPECIFIC PURPOSES, AND REIMBURSEMENTS IN AID" as subheader.
Then table for those fees.
Then "Carried forward" with totals.
I'll need to decide on the amounts. Let's pair them as per the sequence.
First section items (13 items) with 1912 amounts (first 13 numbers after brought forward) and 1913 amounts (next 13 numbers). But we have 25 numbers. Let's assume the first 13 are 1912, the next 12 are 1913, and the 13th 1913 is missing. But the 14th number is 2,000.00, 15th 10.00, 16th 4,222.53, etc. That would make 1913 amounts: 2,000.00, 10.00, 4,222.53, 1,151,311.88, 4,769.55, 1,065,914.90, 4,950.00, 5,051.50, 1,515.00, 1,715.00, 34.00, 44.00. That's 12. Which item gets 2,000.00? Maybe "Money Changers' Licences" 1913 is 4,222.53 (which is 16th). So perhaps the first two numbers (2,000.00 and 10.00) are not part of this section? But they appear right after 500.00. Could be that the 1912 amounts are 13 numbers: 4,110.25; 1,183,200.00; 3,260.00; 1,183,200.00; 92,362.50; 104,240.82; 258.00; 268.00; 20.00; 20.00; 70.00; 70.00; 500.00. Then the 1913 amounts start at 2,000.00? But 2,000.00 for Money Changers? No, Money Changers 1913 is 4,222.53. So maybe the 1913 amounts are: 4,222.53; 1,151,311.88; 4,769.55; 1,065,914.90; 4,950.00; 5,051.50; 1,515.00; 1,715.00; 34.00; 44.00; and then two missing? That's 10. Not 13.
This is too messy. Perhaps the OCR has duplicated some numbers. Notice "1,183,200.00" appears twice in the 1912 list (positions 2 and 4). That might be for Opium Farm and Permits? Both could be same amount. Then 1913 list has "1,151,311.88" and "1,065,914.90". That matches.
So the 1912 list:
Then 1913 list:
But there are two extra numbers before 4,222.53: 2,000.00 and 10.00. Could they be for "Opium Farm" and "Pawnbrokers" 1913? No.
Maybe the 1913 list starts at 2,000.00 for Money Changers? But 2,000 is not 4,222.53. Could be that the 1913 amounts are interleaved with 1912? No.
Given the confusion, I'll assume the OCR has a formatting issue where the two columns are side by side and the OCR read them as a single column. The numbers appear in pairs: 4,110.25 and 4,222.53? But they are far apart.
Let's look at the raw OCR: after "Water-Wheel Licences,...." there is a blank line then "1. FEES OF COURT OR OFFICE..." then "3,338,304,97" then "4,110.25" then "1,183,200.00" then "3,260.00" then "1,183,200.00" then "92,362.50" then "104,240.82" then "258.00" then "268.00" then "20.00" then "20.00" then "70.00" then "70.00" then "500.00" then "2,000.00" then "10,00" then "4,222.53" then "1,151,311.88" then "4,769.55" then "1,065,914.90" then ".4,950.00" then "5,051.50" then "1,515.00" then "1,715.00" then "34,00" then "44.00" then "Analyses,".
Maybe the "3,338,304,97" is the 1913 brought forward? But we already have brought forward line with two numbers. Actually the first line: "Brought forward.$ 3,114,758,99 3,338,304,97". So that's both years. Then the next "3,338,304,97" might be a duplicate or the start of the next column? Could be that the table has two columns: 1912 and 1913, and the OCR read the 1913 column first for the next section? Hmm.
I think the best is to present the data as a single table with all rows, using the descriptions in order and the amounts in order as they appear, but I'll group them by section. Since the user wants proofreading, not data entry, maybe I should just output the text with corrected formatting, preserving the original line breaks but fixing obvious OCR errors (like "Liceuces" -> "Licences", "Fario" -> "Farm", "Chiuese" -> "Chinese", "eatering" -> "entering", "Coutribution" -> "Contribution", "Lensus" -> "Leases", "Au" maybe "An" or just remove). And reconstruct the table by aligning the numbers with descriptions based on the apparent structure.
Given the instruction "Reconstruct tabular data using Markdown table syntax", I must create a table. I'll create two tables: one for Licences, one for Fees.
I'll use the following approach: For the Licences section, I'll list the 13 items and assign the 1912 amounts as the first 13 numbers after the section header (excluding the duplicate brought forward), and the 1913 amounts as the next 13 numbers, but since there are only 12, I'll note the last as missing. However, the numbers 2,000.00 and 10.00 are likely not part of this section. Maybe they are for "Opium Farm" and "Pawnbrokers" 1913? But 2,000 and 10 are too small.
Let's check the numbers after "Analyses,": 10,260.50 and 8,816.50. Those are likely for Analyses 1912 and 1913. Then "Bills of Health," etc. So the Fees section has its own numbers.
Thus, the numbers before "Analyses," belong to the Licences section. There are 26 numbers from 4,110.25 to 44.00. That's exactly 26 numbers. 13 items * 2 = 26. So the first 13 are 1912, the second 13 are 1913. Let's split the list of 26 numbers into two groups of 13.
List of 26 numbers in order:
"4,110.25
1,183,200.00
3,260.00
1,183,200.00
92,362.50
104,240.82
258.00
268.00
20.00
20.00
70.00
70.00
500.00
2,000.00
10,00
4,222.53
1,151,311.88
4,769.55
1,065,914.90
.4,950.00
5,051.50
1,515.00
1,715.00
34,00
44.00"
That's 25 lines. But the 13th is 500.00, 14th 2,000.00, 15th 10.00, 16th 4,222.53, 17th 1,151,311.88, 18th 4,769.55, 19th 1,065,914.90, 20th .4,950.00, 21st 5,051.50, 22nd 1,515.00, 23rd 1,715.00, 24th 34.00, 25th 44.00. That's 25 numbers. We need 26. Perhaps the "3,338,304,97" before "4,110.25" is actually the first 1913 amount for the first item? But that's the brought forward 1913. Or maybe the first item "Money Changers" has only one amount? No.
Maybe the list includes 14 items? Let's count items again: The OCR lists: Money Changers' Liceuces, Opium Fario, Pawnbrokers' Licences, Permits to cut Earth and Stone, New Territories, Pharmacy Ordinance Licences, Pilots' Licences, Prospecting Licences, N.T., Sandalwood Mill Licences, N.T., Special Fruit Licences, Stamp Duties, Steam-Launch Licences, Theatrical Licences, Water-Wheel Licences. That's 14 items? Let's count: 1 Money Changers, 2 Opium, 3 Pawnbrokers, 4 Permits, 5 Pharmacy, 6 Pilots, 7 Prospecting, 8 Sandalwood, 9 Special Fruit, 10 Stamp Duties, 11 Steam-Launch, 12 Theatrical, 13 Water-Wheel. That's 13. Wait, I count 13. But maybe "Opium Fario" and "Opium Farm" are two? No.
Could be that "Pharmacy Ordinance Licences" and "Pilots' Licences" are separate, yes.
If there are 13 items, we need 26 numbers. We have 25. One number missing. Perhaps the "10,00" is actually "10.00" and it's the 1913 amount for "Water-Wheel"? But Water-Wheel 1912 is 500.00, 1913 maybe 10.00? That seems unlikely. Or maybe the "2,000.00" is the 1913 for "Money Changers"? But then 4,222.53 for Opium? No.
Given the difficulty, I'll assume the 1912 amounts are the first 13 numbers (1-13), and the 1913 amounts are the next 13 numbers (14-26), but since we have only 25, the last 1913 amount is missing. I'll insert "..." for the missing.
But the 14th number is 2,000.00, 15th 10.00, 16th 4,222.53, etc. That would assign 2,000.00 to Money Changers 1913, 10.00 to Opium 1913, 4,222.53 to Pawnbrokers 1913, etc. That seems wrong.
Maybe the numbers are paired per item: each item has two numbers adjacent in the OCR? But the OCR shows all numbers in a single column. However, the original table likely had two columns side by side. The OCR might have read the left column (1912) top to bottom, then the right column (1913) top to bottom. That would give all 1912 amounts first, then all 1913 amounts. That matches: first 13 numbers are 1912, next 13 are 1913. But we have 25 numbers, so maybe one 1912 amount is missing or one 1913 amount is missing. Let's assume the first 13 are 1912 (up to 500.00). Then the next 12 are 1913 (2,000.00 to 44.00). The 13th 1913 amount is missing. But then the 1913 amounts would be: Money Changers: 2,000.00? Opium: 10.00? Pawnbrokers: 4,222.53? Permits: 1,151,311.88? Pharmacy: 4,769.55? Pilots: 1,065,914.90? Prospecting: 4,950.00? Sandalwood: 5,051.50? Special Fruit: 1,515.00? Stamp Duties: 1,715.00? Steam-Launch: 34.00? Theatrical: 44.00? Water-Wheel: missing. That seems nonsensical.
Thus, the OCR did not read column by column. It read row by row but the rows are broken across lines. The numbers are interleaved with descriptions? But the descriptions are all listed first, then numbers. So the original might have had the descriptions in the left column, and the two amount columns to the right. The OCR read the left column (descriptions) first, then the two amount columns? But the amounts are not separated by year.
Given the time, I'll produce a proofread text that preserves the original structure as much as possible, with corrected OCR errors, and present the tabular data in a markdown table with the data as it appears, but I'll align the amounts with the descriptions based on the order they appear in the OCR after the section header, assuming the first amount after each description is 1912 and the second is 1913? But the descriptions are not interleaved with amounts.
I think the best is to output the text as a series of paragraphs and then a table for each section, using the data from the OCR but with corrected spelling and formatting. I'll create a table for the Licences section with the 13 items and two columns for amounts, filling in the amounts from the OCR in the order they appear, but I'll note that the alignment is uncertain.
However, the user might be testing my ability to reconstruct tables from OCR. I should do my best.
Let me try to parse the Fees section similarly. After "Analyses," there are numbers: 10,260.50 and 8,816.50. Then descriptions: Bills of Health, Births and Deaths Registration, Cemetery Fees, Certificates to Chinese entering America, Chemists and Druggists Registration, Chinese Cemetery Fees, Companies Registration, Contribution from Chinese Dispensaries, &c. Then numbers: 8,358.50; 10,044.00; 1,711.80; 1,987.10; 931.00; 655.98; 425.00; 225.00; 100.00; 150.00; 2,611.50; 2,794.00; 9,301.30; 9,276.60; 389.94; 210.00. Then description: Contribution from Wei-hai-wei Government for Clerical Assistance, Au (maybe "An" or "Au" is a typo), then numbers 360.00, 430.00. Then Court Fees, 39,679.25, 41,847.77. Crown Leases, 1,820.00, 3,670.00. Deeds Registration, 43,408.95, 51,437.50. Engagement and Discharge of Seamen, 23,784.20, 81,576.20. Carried forward, 5,700,404.19, 5,872,039.39.
This section seems more regular: each description followed by two numbers (1912 and 1913). For example, "Analyses, 10,260.50 8,816.50" -> Analyses: 10,260.50 (1912), 8,816.50 (1913). Then "Bills of Health," but no numbers immediately after? Actually after "Analyses," the next line is "Bills of Health," then "Births and Deaths Registration," then "Cemetery Fees,...." then "Certificates to Chiuese eatering America," then "Chemists and Druggists Registration..................." then "Chinese Cemetery Fees," then "Companies Registration,...................." then "Contribution from Chinese Dispensaries, &c., ............................................................." then numbers. So the numbers for those items are listed after all descriptions. That suggests the OCR read the description column first, then the amount columns. So for the Fees section, there are 9 descriptions (Analyses, Bills of Health, Births and Deaths Registration, Cemetery Fees, Certificates to Chinese entering America, Chemists and Druggists Registration, Chinese Cemetery Fees, Companies Registration, Contribution from Chinese Dispensaries) then 16 numbers? Actually 16 numbers: 8,358.50; 10,044.00; 1,711.80; 1,987.10; 931.00; 655.98; 425.00; 225.00; 100.00; 150.00; 2,611.50; 2,794.00; 9,301.30; 9,276.60; 389.94; 210.00. That's 16 numbers for 9 items? Not matching.
But then there are more descriptions: Contribution from Wei-hai-wei, Court Fees, Crown Leases, Deeds Registration, Engagement and Discharge of Seamen. That's 5 more items, total 14 items. 14 items 2 = 28 numbers. We have 16 + 2 (Wei-hai-wei) + 2 (Court Fees) + 2 (Crown Leases) + 2 (Deeds) + 2 (Engagement) + 2 (Carried forward) = 28 numbers. Yes! The numbers after the first 9 descriptions are actually for all 14 items? Let's see: The first 9 descriptions (Analyses to Contribution from Chinese Dispensaries) are 9 items. Then there are 5 more items (Wei-hai-wei, Court Fees, Crown Leases, Deeds, Engagement). That's 14 items. The numbers list: first 16 numbers? Actually 16 numbers before Wei-hai-wei description. But Wei-hai-wei description appears after those 16 numbers. Then Wei-hai-wei has two numbers (360, 430). Then Court Fees has two numbers (39,679.25, 41,847.77). Crown Leases two numbers (1,820.00, 3,670.00). Deeds two numbers (43,408.95, 51,437.50). Engagement two numbers (23,784.20, 81,576.20). Carried forward two numbers (5,700,404.19, 5,872,039.39). That's 16 + 2+2+2+2+2+2 = 28 numbers. But the first 16 numbers correspond to the first 8 items? 8 items 2 = 16. But we have 9 items before Wei-hai-wei. Analyses is separate with its own numbers (10,260.50, 8,816.50) before the list. So Analyses is item 1 with its numbers given inline. Then the next 8 items (Bills of Health to Contribution from Chinese Dispensaries) have their 16 numbers in the block. Then the remaining 5 items have their numbers inline after each description. That makes sense.
Thus, the Fees section can be reconstructed as:
Carried forward: 5,700,404.19, 5,872,039.39
That works perfectly! So the Fees section is clear.
Now the Licences section: The descriptions are listed first (13 items), then a block of numbers. The block of numbers has 25 numbers. But we need 26. However, note that the first number in the block is "3,338,304,97" which is actually the brought forward 1913? But brought forward already given. Maybe the block starts with the 1912 amounts for the 13 items, then the 1913 amounts for the 13 items. But we have 25 numbers. Let's count the numbers in the block after the section header and before "Analyses,". The OCR shows:
"3,338,304,97
4,110.25
1,183,200.00
3,260.00
1,183,200.00
92,362.50
104,240.82
258.00
268.00
20.00
20.00
70.00
70.00
500.00
2,000.00
10,00
4,222.53
1,151,311.88
4,769.55
1,065,914.90
.4,950.00
5,051.50
1,515.00
1,715.00
34,00
44.00"
That's 26 lines if we count "3,338,304,97" as line 1. But "3,338,304,97" is likely the 1913 brought forward, which is already in the brought forward line. However, it might be the first 1913 amount for the first item? But the brought forward line already has two numbers: 3,114,758.99 and 3,338,304.97. So the "3,338,304,97" appearing again might be a duplicate or the start of the 1913 column for the licences. If we consider that the licences table has its own brought forward? But the text says "Brought forward.$ 3,114,758,99 3,338,304,97" then "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued". So the brought forward is for the whole statement. Then the licences section continues. The numbers after might be the amounts for each licence for 1912 and 1913. But there are 13 items. If the first number "3,338,304,97" is actually the 1913 amount for the first item? That would be weird.
Maybe the OCR has a line "G." then "Brought forward.$ 3,114,758,99 3,338,304,97" then "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued" then the table for licences with two columns: 1912 and 1913. The OCR read the 1912 column first (13 numbers), then the 1913 column (13 numbers). But the 1912 column might start with "4,110.25" and end with "500.00". The 1913 column might start with "2,000.00" and end with "44.00". But then there is an extra "3,338,304,97" before "4,110.25". That could be the 1913 brought forward repeated. Or it could be the 1912 amount for the first item? But 3,338,304.97 is huge, not plausible for Money Changers.
Thus, I think the "3,338,304,97" is a stray duplicate. The actual 1912 amounts are the 13 numbers from 4,110.25 to 500.00. The 1913 amounts are the 13 numbers from 2,000.00 to 44.00? But that gives 13 numbers: 2,000.00, 10.00, 4,222.53, 1,151,311.88, 4,769.55, 1,065,914.90, 4,950.00, 5,051.50, 1,515.00, 1,715.00, 34.00, 44.00. That's only 12. Wait, count: 1)2,000.00 2)10.00 3)4,222.53 4)1,151,311.88 5)4,769.55 6)1,065,914.90 7)4,950.00 8)5,051.50 9)1,515.00 10)1,715.00 11)34.00 12)44.00. That's 12. So one missing. Perhaps the "3,338,304,97" is the first 1913 amount? But that's the brought forward. Or maybe the 1913 amounts start at "4,222.53" and the "2,000.00" and "10.00" are something else (like totals for something). But they appear in the sequence.
Given the Fees section worked out perfectly with the pattern: descriptions listed first, then a block of numbers for the first group, then descriptions with inline numbers for the rest. For Licences, all 13 items are listed first, then a block of numbers. That block should contain 26 numbers (13*2). We have 25 numbers after the duplicate. If we ignore the duplicate "3,338,304,97", we have 25 numbers. One number missing. Could be that the last number for Water-Wheel 1913 is missing. The last number in the block is 44.00. That would be the 12th 1913 amount. So the 13th 1913 amount is missing. I'll assume it's missing and put "...".
But which amount corresponds to which item? If the block is 1912 amounts first (13 numbers), then 1913 amounts (13 numbers), then the 1912 amounts are: 4,110.25; 1,183,200.00; 3,260.00; 1,183,200.00; 92,362.50; 104,240.82; 258.00; 268.00; 20.00; 20.00; 70.00; 70.00; 500.00. The 1913 amounts would be the next 13 numbers: but we have only 12 numbers after that: 2,000.00; 10.00; 4,222.53; 1,151,311.88; 4,769.55; 1,065,914.90; 4,950.00; 5,051.50; 1,515.00; 1,715.00; 34.00; 44.00. That's 12. So perhaps the 1913 amounts start at 4,222.53? And the 2,000.00 and 10.00 are not part of this? But they are in the sequence.
Maybe the block includes both years interleaved per item? But the Fees section had a block of 16 numbers for 8 items (two per item). That block was 1912 and 1913 interleaved? Actually the Fees block: 8,358.50; 10,044.00; 1,711.80; 1,987.10; 931.00; 655.98; 425.00; 225.00; 100.00; 150.00; 2,611.50; 2,794.00; 9,301.30; 9,276.60; 389.94; 210.00. That's 16 numbers for 8 items. If they are interleaved per item (1912, 1913, 1912, 1913...), then item1: 8,358.50 (1912), 10,044.00 (1913); item2: 1,711.80 (1912), 1,987.10 (1913); item3: 931.00 (1912), 655.98 (1913); item4: 425.00 (1912), 225.00 (1913); item5: 100.00 (1912), 150.00 (1913); item6: 2,611.50 (1912), 2,794.00 (1913); item7: 9,301.30 (1912), 9,276.60 (1913); item8: 389.94 (1912), 210.00 (1913). That matches the order of descriptions: Bills of Health, Births and Deaths Registration, Cemetery Fees, Certificates to Chinese entering America, Chemists and Druggists Registration, Chinese Cemetery Fees, Companies Registration, Contribution from Chinese Dispensaries. Perfect! So the block is interleaved per item: 1912 then 1913 for each item in order.
Thus, for the Licences section, the block of numbers after the descriptions should be interleaved per item: 1912, 1913 for each of the 13 items. That would be 26 numbers. The block starts with "3,338,304,97" then 25 more numbers. If we consider the first number "3,338,304,97" as the 1912 amount for the first item? But that's huge. Maybe the first item is "Brought forward" but that's already done. The licences section might have a "Brought forward" line again? The text says "Brought forward.$ 3,114,758,99 3,338,304,97" then "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued". So the brought forward is before the licences section. The licences section might start with its own first item. The block of numbers might be for the licences items only. But the first number in the block is "3,338,304,97" which is the same as the 1913 brought forward. Could be a scanning artifact.
If we skip that, the next numbers are 4,110.25; 1,183,200.00; 3,260.00; 1,183,200.00; 92,362.50; 104,240.82; 258.00; 268.00; 20.00; 20.00; 70.00; 70.00; 500.00; 2,000.00; 10.00; 4,222.53; 1,151,311.88; 4,769.55; 1,065,914.90; 4,950.00; 5,051.50; 1,515.00; 1,715.00; 34.00; 44.00. That's 25 numbers. For 13 items interleaved, we need 26. So one missing. The pattern would be: item1: 4,110.25 (1912), 1,183,200.00 (1913)? But that would assign Opium Farm 1913 as 1,183,200.00? Actually the second number is 1,183,200.00. If interleaved, item1 (Money Changers): 4,110.25 (1912), 1,183,200.00 (1913) - but 1,183,200 is huge for Money Changers. So not interleaved per item.
Maybe the block is all 1912 amounts first, then all 1913 amounts. That gave 13 and 12. The Fees section block was interleaved because the descriptions were in a different order? Actually the Fees section had the first 8 descriptions (after Analyses) listed, then a block of 16 numbers interleaved. That suggests the OCR read the table row by row for those 8 items: description, 1912, 1913. But the OCR output separated descriptions and numbers. For Licences, the OCR listed all descriptions first, then all numbers. The numbers might be in row-major order: 1912 for item1, 1913 for item1, 1912 for item2, 1913 for item2, etc. But the numbers don't match that.
Let's test interleaved on the Licences numbers (excluding the first duplicate). Sequence:
If interleaved per item (13 items), we need pairs: (1,2), (3,4), (5,6), (7,8), (9,10), (11,12), (13,14), (15,16), (17,18), (19,20), (21,22), (23,24), (25,26). But we have 25 numbers. So pairs would be:
Item1: 4,110.25, 1,183,200.00
Item2: 3,260.00, 1,183,200.00
Item3: 92,362.50, 104,240.82
Item4: 258.00, 268.00
Item5: 20.00, 20.00
Item6: 70.00, 70.00
Item7: 500.00, 2,000.00
Item8: 10.00, 4,222.53
Item9: 1,151,311.88, 4,769.55
Item10: 1,065,914.90, 4,950.00
Item11: 5,051.50, 1,515.00
Item12: 1,715.00, 34.00
Item13: 44.00, (missing)
This seems random. Not plausible.
Given the Fees section worked with interleaved block for the first 8 items, and the Licences section has 13 items, perhaps the Licences block is also interleaved but the OCR missed the first 1912 amount for the first item? The first number in the block is "3,338,304,97" which might be the 1912 amount for Money Changers? No.
Wait, the Fees section had the first item (Analyses) with numbers inline, then the next 8 items had a block of 16 numbers interleaved. The Licences section has no item with inline numbers; all 13 items are listed then a block. The block might be 26 numbers interleaved. The OCR shows 26 lines if we include "3,338,304,97". Let's count lines in the OCR from "3,338,304,97" to "44.00": there are 26 lines. So 26 numbers. If they are interleaved per item, then the 13 items correspond to 13 pairs. The first pair would be (3,338,304.97, 4,110.25). That would be Money Changers: 1912=3,338,304.97, 1913=4,110.25. But 3,338,304.97 is the brought forward total, not a licence amount. So maybe the block includes the brought forward again? The text says "Brought forward.$ 3,114,758,99 3,338,304,97" then "LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED,—Continued". The "Continued" suggests the licences section continues from a previous page. The brought forward might be the cumulative total up to that point. Then the licences section lists individual licences. The block of numbers might be the amounts for each licence for 1912 and 1913. But the first number in the block is the 1913 brought forward? That doesn't make sense.
Perhaps the OCR has a formatting issue: the "3,338,304,97" is actually the 1913 amount for the last item of the previous section? But the previous section is just the brought forward.
Given the time, I will produce a proofread version that corrects obvious OCR errors and presents the data in a clear markdown format, with tables for the two sections. For the Licences section, I will create a table with the 13 items and two columns for amounts, using the numbers that seem most plausible based on typical revenue figures. But I must not invent data. However, the instruction says "If a word or phrase is clearly missing due to OCR damage, insert ... in its place." So for missing amounts, I can put "...".
I'll assume the 1912 amounts are the first 13 numbers after the section header (excluding the duplicate brought forward), and the 1913 amounts are the next 13 numbers, with the last missing. But the 1913 amounts would be 2,000.00, 10.00, 4,222.53, 1,151,311.88, 4,769.55, 1,065,914.90, 4,950.00, 5,051.50, 1,515.00, 1,715.00, 34.00, 44.00, ... That seems odd but I'll present as is.
Alternatively, I could present the Licences section as a table with the items and the amounts as they appear in the OCR in two columns, but I'll note the uncertainty. However, the user wants a clean proofread.
Given the Fees section is clear, I'll focus on that. For the Licences section, I'll replicate the OCR structure but corrected.
Let's write the markdown.
First, the header: "(08)" maybe a page number. Then "Specify each separate Tax or Duty according to the classification employed in the Estimates," then "COMPARATIVE YEARLY STATEMENT" then "Amount collected in the Year 1912 | Amount collected in the Year 1913" then "G." then "Brought forward. $3,114,758.99 $3,338,304.97"
Then "## LICENCES AND INTERNAL REVENUE NOT OTHERWISE SPECIFIED—Continued"
Then a table:
| Tax/Duty | 1912 | 1913 |
|---|---|---|
| Money Changers' Licences | 4,110.25 | 4,222.53 |
| Opium Farm | 1,183,200.00 | 1,151,311.88 |
| Pawnbrokers' Licences | 3,260.00 | 4,769.55 |
| Permits to cut Earth and Stone, New Territories | 1,183,200.00 | 1,065,914.90 |
| Pharmacy Ordinance Licences | 92,362.50 | 4,950.00 |
| Pilots' Licences | 104,240.82 | 5,051.50 |
| Prospecting Licences, N.T. | 258.00 | 1,515.00 |
| Sandalwood Mill Licences, N.T. | 268.00 | 1,715.00 |
| Special Fruit Licences | 20.00 | 34.00 |
| Stamp Duties | 20.00 | 44.00 |
| Steam-Launch Licences | 70.00 | ... |
| Theatrical Licences | 70.00 | ... |
| Water-Wheel Licences | 500.00 | ... |
But I need to decide which 1913 amount goes with which. The plausible pairing based on magnitude: Opium Farm 1913 ~1.15M, Permits ~1.06M, Pawnbrokers ~4,769, Pharmacy ~4,950, Pilots ~5,051, Prospecting ~1,515, Sandalwood ~1,715, Special Fruit ~34, Stamp Duties ~44. Money Changers ~4,222. Steam-Launch, Theatrical, Water-Wheel missing. The numbers 2,000 and 10 are extra. Maybe they are for Steam-Launch and Theatrical? 2,000 for Steam-Launch? 10 for Theatrical? But 70 for 1912. Could be.
I'll pair them in order of the 1913 numbers after the first 13 1912 numbers. The 1913 numbers in the OCR after the 13th 1912 number (500.00) are: 2,000.00, 10.00, 4,222.53, 1
(08)
Specify each separate Tax or Duty according to the classification employed in the Estimates,
COMPARATIVE YEARLY STATE-
Amount collected Amount collected in the Year 1912. |in the Year 1913.
G.
Brought forward.$
3,114,758,99
LICENCES AND INTERNAL REVENUE NOT OTHERWISE
SPECIFIED,—Continued,
Money Changers' Liceuces,
Opium Fario,
Pawnbrokers' Licences,..
Permits to cut Earth and Stone, New Territories,.
Pharmacy Ordinance Licences, - • • •
Pilots' Licences,.
Prospecting Licences, N.T.,.......................
Sandalwood Mill Licences, N.T.,
Special Fruit Licences,....
Stamp Duties,.........
Steam-Launch Licences,
Theatrical Licences,
Water-Wheel Licences,....
PURPOSES, AND REIMBURSEMENTS IN AID :-
3,338,304,97
4,110.25
1,183,200.00
3,260.00
1,183,200.00
92,362.50
104,240.82
258.00
268.00
20.00
20.00
70.00
70.00
500.00
2,000.00
10,00
4,222.53
1,151,311.88
4,769.55
1,065,914.90
.4,950.00
5,051.50
1,515.00
1,715.00
34,00
44.00
Analyses,
10,260.50
8,816.50
Bills of Health,
Births and Deaths Registration,
Cemetery Fees,....
Certificates to Chiuese eatering America,
Chemists and Druggists Registration...................
Chinese Cemetery Fees,
Companies Registration,....................
Contribution from Chinese Dispensaries, &c., .............................................................
8,358.50
10,044.00
1,711,80
1,987.10
931.00
655,98
425,00
225.00
100.00
150,00
2,611.50
2,794.00
9,301.30
9,276.60
389,94
210.00
·Coutribution from Wei-hai-wei Government for Clerical Assistance,
Au
360.00
430.00
Court Fees,
39,679.25
41,847.77
Crown Lensus,
1.820.00
3,670.00
Deeds Registration,
43,408.95
51,437.50
Engagement and Discharge of Seamen,
23,784.20
81,576.20
Carried forward,...........................................
5,700,404.19
5,872,039.39
No comments yet.
Private notes are available after approval.