The user wants me to proofread OCR output of a historical table about vessels cleared at ports. This appears to be a statistical table showing vessels cleared from Hong Kong ports to various countries, with columns for British and Foreign vessels, with cargoes and in ballast.
Let me analyze the OCR text and reconstruct the table properly. The OCR is quite messy with broken lines, misaligned columns, and garbled numbers.
First, let me understand the structure:
Looking at the data rows, I need to parse each country and its statistics.
Let me reconstruct row by row:
The column structure seems to be:
Wait, the header line: "Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews."
Let me count: That's 10 columns after Countries? Or maybe:
But the header only shows 10 "Crews/Vessels/Tons" entries. Let me recount: "Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." That's 10 items. But each category needs 3 columns (Vessels, Tons, Crews).
Actually, the header row in the OCR: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." Wait, the OCR shows: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." but it's split across lines.
Looking at the original: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." - that's 12 columns.
But the OCR text shows: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." No, it shows: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." Actually the text says: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." Wait, I see: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." but in the OCR it's broken.
Let me look at the OCR header lines:
"WITH CARGOES.
Vessels. Tons.
BRITISH.
IN BALLAST.
TOT.L..
FOREIGN.
WITH CARGOES,
Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews."
This is messy. Let me interpret the intended columns based on typical Hong Kong Blue Book statistical tables.
Typical structure for "Vessels Cleared" tables:
That would be 1 + 33 + 33 = 19 columns. But the header only shows 10 sub-headers.
Alternatively, maybe it's:
But the header "TOT.L.." suggests Total for British, and "FOREIGN. WITH CARGOES," suggests Foreign With Cargoes, but then what about Foreign In Ballast and Foreign Total?
Looking at the data rows, for "Coast of China and Formosa" we see: "1,352 1,387,028 58,632 49 40,831 1,391 1,401 1,427,859 60,023 17,593 1,595,080 237,583"
That's 12 numbers. Let me parse:
1,352 | 1,387,028 | 58,632 | 49 | 40,831 | 1,391 | 1,401 | 1,427,859 | 60,023 | 17,593 | 1,595,080 | 237,583
If British With Cargoes: 1,352 vessels, 1,387,028 tons, 58,632 crews
British In Ballast: 49 vessels, 40,831 tons, 1,391 crews
British Total: 1,401 vessels, 1,427,859 tons, 60,023 crews
Foreign With Cargoes: 17,593 vessels, 1,595,080 tons, 237,583 crews? But that's only 3 numbers for foreign, missing In Ballast and Total.
Wait, 12 numbers = 4 groups of 3. So maybe:
But then where is Foreign In Ballast and Foreign Total? Maybe the table only shows Foreign With Cargoes? But the header says "FOREIGN. WITH CARGOES," and then "Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." - that's 10 items, not 12.
Let me check another row: "Cochin China" - "48 34,971 1,348 63 53,439 1,905 111 88,410 3,253 53 36,023 1,334"
That's 12 numbers: 48, 34971, 1348, 63, 53439, 1905, 111, 88410, 3253, 53, 36023, 1334
Groups of 3:
So 4 groups of 3. For Cochin China, British With Cargoes: 48 vessels, 34,971 tons, 1,348 crews
British In Ballast: 63 vessels, 53,439 tons, 1,905 crews
British Total: 111 vessels, 88,410 tons, 3,253 crews
Foreign With Cargoes: 53 vessels, 36,023 tons, 1,334 crews
But then Foreign In Ballast and Foreign Total are missing? Or maybe the table only has Foreign With Cargoes? But the header mentions "FOREIGN. WITH CARGOES," and then the column headers continue.
Wait, the header line: "Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." That's 10 column headers. But we have 12 data columns per row (for rows that have full data).
Let me count the header items more carefully from the OCR:
"Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." - that would be 12 if complete. But the OCR shows: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." Actually in the text it's: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." No, the OCR text says: "Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." Wait, I need to look at the raw OCR:
"Vessels. Tons.
BRITISH.
IN BALLAST.
TOT.L..
FOREIGN.
WITH CARGOES,
Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews."
This is very garbled. The "Vessels. Tons." at top might be for "WITH CARGOES" under British. Then "BRITISH." is a section header. "IN BALLAST." is next section. "TOT.L.." is Total for British. "FOREIGN." is section. "WITH CARGOES," is subsection. Then the column headers: "Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews."
But that's only 10 column headers for what appears to be 12 data columns. Unless the first "Vessels. Tons." are the first two columns of the first group, and "Crews" is the third, but it's listed later.
Actually, typical layout: The column headers run across the top. The OCR has read them in reading order (left to right, top to bottom) but the original had multi-row headers.
Let me think of the original table structure:
| Countries | British | Foreign | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| With Cargoes | In Ballast | Total | With Cargoes | In Ballast | Total | |||||||||||||
| Vessels | Tons | Crews | Vessels | Tons | Crews | Vessels | Tons | Crews | Vessels | Tons | Crews | Vessels | Tons | Crews | Vessels | Tons | Crews |
That's 19 columns. But the data only shows 12 numbers per row for full rows. So maybe the table only has British (With Cargoes, In Ballast, Total) and Foreign (With Cargoes only)? But that seems odd.
Wait, look at the totals row: "Totally 2,304 2,200.074 209,074 99,770 180 181,618 6,017 2,484 2,360,689 104.67 18,960 2,094,214 29,3T."
Let me parse totals: 2,304 | 2,200,074 | 209,074 | 99,770 | 180 | 181,618 | 6,017 | 2,484 | 2,360,689 | 104,67? | 18,960 | 2,094,214 | 29,3T
That's 13 numbers? Let me count: 2304, 2200074, 209074, 99770, 180, 181618, 6017, 2484, 2360689, 10467, 18960, 2094214, 293T (maybe 293,770?). 13 numbers.
But other rows have 12. Hmm.
Let me check "India & Singapore": "139 176,480 7,278 14 13,298 326 153 189,778 7,004 31 36,304 1,228" - 12 numbers.
"Japan": "61 78,335 3,821 13 11,500 397 74 89,895 4,218 56 59,024 4,206" - 12 numbers.
"Java and other Islands": "1 510 19 3 2,849 3,359 96 8,504 670" - only 9 numbers? "1 510 19 3 2,849 3,359 96 8,504 670" - that's 9. But maybe missing foreign.
"Macao": "377 214,073 13,7309 176 25 379 214,249 13,764 974 1-45,258 24,652" - let's parse: 377, 214073, 137309? 176, 25, 379, 214249, 13764, 974, 145258, 24652 - that's 11 numbers? 377, 214073, 137309 (3), 176, 25, 379 (3), 214249, 13764, 974 (3), 145258, 24652 (2) = 11. Not consistent.
This is very messy. Given the OCR quality, I need to reconstruct the table as best as possible, aligning columns based on the header description and typical format.
The title: "No. 2.-NUMBER, TONNAGE, and CREWS of VESSELS CLEARED at PORTS" - likely "at Ports in Hong Kong" or similar.
The header mentions: COUNTRIES TO WHICH DEPARTED, WITH CARGOES, BRITISH, IN BALLAST, TOTAL, FOREIGN, WITH CARGOES.
Given the data, I think the table has these columns:
But the data only fills up to column 13 for most rows. The totals row has more numbers.
Wait, the totals row: "2,304 2,200.074 209,074 99,770 180 181,618 6,017 2,484 2,360,689 104.67 18,960 2,094,214 29,3T."
If we map:
That makes sense! So the columns are:
But the header says "FOREIGN. WITH CARGOES," implying there might be Foreign In Ballast and Foreign Total as well. However, the data doesn't show them for individual countries. Maybe the table only shows Foreign With Cargoes, and the totals row only shows Foreign With Cargoes totals.
But then why does the header have "Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews." - 10 items? That would correspond to:
Total 12. But the header shows 10. Unless the first "Vessels. Tons." are separate.
Let me re-read the OCR header lines carefully:
"No. 2.-NUMBER, TONNAGE, and CREWS of VESSELS CLEARED at PORTS
COUNTRIES
TO WHICH DEPARTED.
WITH CARGOES.
Vessels. Tons.
BRITISH.
IN BALLAST.
TOT.L..
FOREIGN.
WITH CARGOES,
Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews."
This is the OCR reading order. The original table likely had a multi-row header:
Row 1: No. 2.-NUMBER, TONNAGE, and CREWS of VESSELS CLEARED at PORTS [maybe "IN HONG KONG" or similar]
Row 2: COUNTRIES TO WHICH DEPARTED | WITH CARGOES | BRITISH | IN BALLAST | TOTAL | FOREIGN | WITH CARGOES | ...
Row 3: | Vessels | Tons | Crews | Vessels | Tons | Crews | Vessels | Tons | Crews | Vessels | Tons | Crews | ...
But the OCR has jumbled the header rows.
Given the data, I'll assume the table has 13 columns (Country + 12 data columns):
And the totals row confirms this: 12 data numbers.
For rows that have fewer numbers (like Java, Macao), some columns are blank or zero.
Now let me parse each country row carefully.
I'll go through the OCR text line by line and extract data.
The OCR text after header:
"Australia and
Avetr Ze land...}
British North
America,
British North
28
30,145 1,209|
:
:
}
7
8,578 248
:
28
30,145 1,209)||
7
8,578 248
14
16,531 207
132
17
Borneo,
1,065 95
:
3
1,065
95
Coast of China
and Formosa...-
1,352 1,387,028 58,632
49 40,831 1,391 1,401 1,427,859 60,023 17,593 1,595,080|237,583|
Cochin China,......
48
34,971 1,348
63 53,439 1,905
111
88,410 3,253 53
36,023|| 1,334
Continent of
Europe,
30
51,590 3,475
30
51,590 3,475
391
70,818 4,753
:
Great Britain,............
33
49,504 1,392
33
49,564 1,392
4,417) 131
India & Singapore,
139
176,480 7,278 14 13,298 326
153
189,778 7,004
31
36,304 1,228
Japan,
61
78,335 3,821
13 11,500 397
74
89,895 4,218 56
59,024 4,206|
Java and other
Islands in the
1
510
Indian Archi-
19 3
2,849
3,359 96
8,504
670
pelago,
Macao,....
377
214,073 13,7309
to
176
25
379 214,249 13,764 974 1-45,258 24,652|
Mauritius,
1
300 38
:
1
:
1,300
38
:
North Pacific,
U
8 8,400 280
280
Pellew Islands,
1
290
11
:
1
290
11
:
:
:
÷
Philippine Islands, 26 12,470 900 16 11,388 269
41 23,858 1,169
33
12,899 1,083
Ports in Hainan
and the Gulf of Tonquin,.......
91 34,097 2,566
3,660 100
96
37,757 2,720
32,975 1,987
Russia in Asin,
:
***
:
***
:
1.200
173
Sandwich Islands,
:
..
904
315
Siam,.......
51 35,082) 1,648
7 5,924 187
58
41,006 1,735
23
12.125
509!
South Africa,
A
South America,
South Pacific,
United States of
America,
56
93,496 3,451
:
:
:
:
:
:
498
181
+44
:
19
56 93,496 3,45]||
38.242
Totally
2,304 2,200.074
209,074 99,770 180 181,618 6,017|2,484 2,360,689 104.67|| 18,9602,094,214 29,3T."
This is extremely messy. I need to reconstruct each row.
Let me start with "Australia and New Zealand". The OCR shows "Australia and Avetr Ze land...}" then "British North America," then "British North" with numbers. It seems "Australia and New Zealand" might have no data? Or the data is for "British North America"? Actually, "British North America" appears as a country. "Australia and New Zealand" might be a country with no vessels? But then "British North" appears with 28, 30,145, 1,209 etc.
Look at the structure: "Australia and Avetr Ze land...}" then "British North America," then "British North" and numbers. The "}" might indicate a bracket grouping. Perhaps "Australia and New Zealand" and "British North America" are grouped? But the numbers follow "British North".
Let me read: "Australia and Avetr Ze land...} British North America, British North 28 30,145 1,209| : : } 7 8,578 248 : 28 30,145 1,209)|| 7 8,578 248 14 16,531 207 132 17"
This is confusing. Perhaps the OCR has merged two rows. "Australia and New Zealand" might have data: 28, 30,145, 1,209 for British With Cargoes? Then 7, 8,578, 248 for British In Ballast? Then 28, 30,145, 1,209 repeated? Then 7, 8,578, 248 repeated? Then 14, 16,531, 207 for British Total? Then 132, 17 for Foreign? But 132, 17 is only two numbers.
Wait, "132 17" then "Borneo," so maybe Foreign With Cargoes: 132 vessels, 17? No, tons and crews missing.
Let me look at the pattern for other countries. For "Coast of China and Formosa", we have 12 numbers. For "Cochin China", 12 numbers. For "Continent of Europe", we have: "30 51,590 3,475 30 51,590 3,475 391 70,818 4,753" - that's 9 numbers, then ":" and "Great Britain". So Continent of Europe has only 9 numbers? 30, 51590, 3475 (British With Cargoes), 30, 51590, 3475 (British In Ballast? But same as With Cargoes?), 391, 70818, 4753 (British Total? But 391 != 30+30). Actually 30+30=60, not 391. So maybe the second group is Foreign With Cargoes? 30, 51590, 3475 for Foreign With Cargoes? Then British Total: 391, 70818, 4753? But British Total should be sum of British With Cargoes and In Ballast. If British In Ballast is missing (zeros), then British Total = British With Cargoes = 30, 51590, 3475. But 391 is different.
Perhaps for Continent of Europe, the data is: British With Cargoes: 30, 51590, 3475; British In Ballast: 0,0,0 (not shown); British Total: 30, 51590, 3475; Foreign With Cargoes: 391, 70818, 4753. That would be 12 numbers: 30,51590,3475, 0,0,0, 30,51590,3475, 391,70818,4753. But OCR shows only 9 numbers: "30 51,590 3,475 30 51,590 3,475 391 70,818 4,753". That's three groups of 3. So maybe the table only has three groups: British With Cargoes, British Total, Foreign With Cargoes? But then what about British In Ballast?
Look at "Great Britain": "33 49,504 1,392 33 49,564 1,392 4,417) 131" - 33,49504,1392; 33,49564,1392; 4417,131? That's 8 numbers. 4,417) 131 might be 4,417 and 131? But need 3 numbers for third group.
"India & Singapore": "139 176,480 7,278 14 13,298 326 153 189,778 7,004 31 36,304 1,228" - 12 numbers. Good.
"Japan": "61 78,335 3,821 13 11,500 397 74 89,895 4,218 56 59,024 4,206" - 12 numbers.
"Java and other Islands in the Indian Archipelago": "1 510 19 3 2,849 3,359 96 8,504 670" - 9 numbers.
"Macao": "377 214,073 13,7309 176 25 379 214,249 13,764 974 1-45,258 24,652" - let's count: 377, 214073, 137309 (3), 176, 25, 379 (3), 214249, 13764, 974 (3), 145258, 24652 (2) = 11. The last group has only 2 numbers.
"Mauritius": "1 300 38 1 1,300 38" - 6 numbers? "1 300 38 : 1 : 1,300 38" - maybe 1,300,38 for British With Cargoes? Then 1,300,38 for British Total? Foreign missing.
"North Pacific": "8 8,400 280 8. 8,490 280" - 6 numbers.
"Pellew Islands": "1 290 11 1 290 11" - 6 numbers.
"Philippine Islands": "26 12,470 900 16 11,388 269 41 23,858 1,169 33 12,899 1,083" - 12 numbers.
"Ports in Hainan and the Gulf of Tonquin": "91 34,097 2,566 3,660 100 96 37,757 2,720 32,975 1,987" - 10 numbers? 91,34097,2566 (3), 3660,100,96? 3660,100,96 (3), 37757,2720 (2), 32975,1987 (2) = 10. Not 12.
"Russia in Asia": "1.200 173" - 2 numbers? Only Foreign With Cargoes?
"Sandwich Islands": "904 315" - 2 numbers.
"Siam": "51 35,082 1,648 7 5,924 187 58 41,006 1,735 23 12,125 509" - 12 numbers.
"South Africa": "A" - no data?
"South America": blank
"South Pacific": blank
"United States of America": "56 93,496 3,451 498 181 +44 19 56 93,496 3,45 38.242" - messy.
"Totally": "2,304 2,200.074 209,074 99,770 180 181,618 6,017 2,484 2,360,689 104.67 18,960 2,094,214 29,3T" - 13 numbers? Let's parse as: 2304, 2200074, 209074, 99770, 180, 181618, 6017, 2484, 2360689, 104670, 18960, 2094214, 293770. That's 13. But 13 numbers for 12 columns? Maybe the first number is something else? Or "Totally" is not a country but a label, and the numbers start from British With Cargoes Vessels.
Actually, 12 data columns:
Totals row should have 12 numbers. But we have 13. Let's see: 2304 (Br WC Ves), 2200074 (Br WC Tons), 209074 (Br WC Crews), 99770 (Br IB Ves? but 99,770 vessels seems high), 180 (Br IB Ves?), 181618 (Br IB Tons), 6017 (Br IB Crews), 2484 (Br Tot Ves), 2360689 (Br Tot Tons), 104670 (Br Tot Crews), 18960 (For WC Ves), 2094214 (For WC Tons), 293770 (For WC Crews). That's 13. If Br IB Ves is 180, then 99770 is not Br IB Ves. Maybe 99770 is Br IB Tons? But then Br IB Ves is missing. The sequence: after Br WC Crews (209074), next is 99770. If columns are Br WC V, Br WC T, Br WC C, Br IB V, Br IB T, Br IB C, Br Tot V, Br Tot T, Br Tot C, For WC V, For WC T, For WC C. Then:
Doesn't align.
Alternative: The table has Br WC, Br IB, Br Tot, For WC, For IB, For Tot. That's 6 groups * 3 = 18 columns. But data rows have 12 numbers. So not.
Given the difficulty, I think the best approach is to reconstruct the table as it appears in the OCR, cleaning up obvious errors, and presenting it in a markdown table with the columns as they can be discerned. Since the user wants proofreading of OCR output, I should output the corrected text with the table reconstructed.
I'll define the columns based on the header description: COUNTRIES TO WHICH DEPARTED, then for BRITISH: WITH CARGOES (Vessels, Tons, Crews), IN BALLAST (Vessels, Tons, Crews), TOTAL (Vessels, Tons, Crews), for FOREIGN: WITH CARGOES (Vessels, Tons, Crews). That's 13 columns including country.
I'll parse each country row to the OCR text sequentially, grouping numbers into sets of 3 for each category.
Let me go through the OCR text linearly and assign.
The text after "COUNTRIES TO WHICH DEPARTED." and header lines:
"Australia and
Avetr Ze land...}
British North
America,
British North
28
30,145 1,209|
:
:
}
7
8,578 248
:
28
30,145 1,209)||
7
8,578 248
14
16,531 207
132
17
Borneo,
1,065 95
:
3
1,065
95"
It seems "Australia and New Zealand" and "British North America" might be two separate entries. The "}" and "{" might indicate a bracket for a group. But the numbers appear after "British North". Let's assume the first country is "Australia and New Zealand" with no data? Or the data for "Australia and New Zealand" is the first set: 28, 30,145, 1,209? But then "British North America" follows.
Actually, looking at the pattern: "Australia and New Zealand...} British North America, British North 28 30,145 1,209". The "British North" might be a continuation of "British North America". In some tables, "British North America" is a country. "Australia and New Zealand" is another. The OCR might have lost a line break.
Let me check the totals: In totals, British With Cargoes vessels = 2,304. If Australia and New Zealand had 28, that's plausible.
I'll treat "Australia and New Zealand" as a country with data: British With Cargoes: 28 vessels, 30,145 tons, 1,209 crews. Then British In Ballast: 7 vessels, 8,578 tons, 248 crews. Then British Total: 28? Wait, the OCR shows "28 30,145 1,209)|| 7 8,578 248 14 16,531 207". That suggests British Total: 14 vessels, 16,531 tons, 207 crews? But 28+7=35, not 14. So maybe the first 28,30145,1209 is for British With Cargoes, then 7,8578,248 for British In Ballast, then 14,16531,207 for Foreign With Cargoes? But 14,16531,207 for Foreign. Then "132 17" might be for next country? But "Borneo" follows.
"132 17 Borneo," - maybe Foreign In Ballast? 132 vessels, 17? No.
Let's look at "Borneo": "Borneo, 1,065 95 : 3 1,065 95". That's 6 numbers: 1,065, 95, 3, 1,065, 95. Could be British With Cargoes: 3 vessels, 1,065 tons, 95 crews? But 1,065 and 95 appear twice. Maybe British With Cargoes: 1,065 tons, 95 crews? But vessels missing. "1,065 95" then ": 3 1,065 95" - perhaps vessels=3, tons=1,065, crews=95 for British With Cargoes. Then British In Ballast: none (colon). Then British Total: same. Then Foreign With Cargoes: none.
But the pattern for other countries: first three numbers are British With Cargoes (V, T, C). For Borneo, if first group is "1,065 95" that's only two numbers. Then "3 1,065 95" is three numbers. So likely British With Cargoes: 3 vessels, 1,065 tons, 95 crews. The preceding "1,065 95" might be a duplicate or misread.
Similarly, "Coast of China and Formosa" has 12 numbers clear.
"Cochin China" has 12 numbers.
"Continent of Europe" has 9 numbers: 30, 51590, 3475, 30, 51590, 3475, 391, 70818, 4753. That's three groups. Could be British With Cargoes, British Total, Foreign With Cargoes. But British In Ballast missing (assumed zero). British Total = British With Cargoes (30,51590,3475) but then third group 391,70818,4753 is Foreign With Cargoes. But why is British Total repeated? Actually the second group is same as first. So maybe the table has only British With Cargoes, British Total, Foreign With Cargoes. But then what about British In Ballast? For Coast of China, we have 4 groups: British WC, British IB, British Tot, Foreign WC. For Continent of Europe, British IB might be zero and not shown? But the OCR shows two identical groups for British WC and British Tot? Or the second group is British IB with same numbers? Unlikely.
Let's check "Great Britain": "33 49,504 1,392 33 49,564 1,392 4,417) 131". Three groups: 33,49504,1392; 33,49564,1392; 4417,131. The third group has only two numbers. 4,417 and 131. Could be Foreign With Cargoes: 4,417 vessels, 131? tons? crews? Missing one.
"India & Singapore": 12 numbers, four groups.
"Japan": 12 numbers, four groups.
"Java": 9 numbers, three groups.
"Macao": 11 numbers, groups of 3,3,3,2.
"Mauritius": 6 numbers, two groups.
"North Pacific": 6 numbers, two groups.
"Pellew Islands": 6 numbers, two groups.
"Philippine Islands": 12 numbers, four groups.
"Ports in Hainan": 10 numbers.
"Russia in Asia": 2 numbers.
"Sandwich Islands": 2 numbers.
"Siam": 12 numbers, four groups.
"South Africa": "A" - maybe no data.
"South America": blank.
"South Pacific": blank.
"United States": messy.
"Totally": 13 numbers.
Given the inconsistency, I think the original table likely had a fixed number of columns, and the OCR has dropped some numbers or merged lines. As a proofreader, I should reconstruct the table with the columns as defined by the header, and fill in the data as best as possible from the OCR, marking missing with ... or leaving blank.
The header says: "COUNTRIES TO WHICH DEPARTED. WITH CARGOES. Vessels. Tons. BRITISH. IN BALLAST. TOT.L.. FOREIGN. WITH CARGOES, Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews."
This suggests the columns are:
Let's assume the table has 13 columns (Country + 12 data). I'll create a markdown table with these 13 columns.
I'll parse each country from the OCR text in order.
List of countries in order as they appear:
Now extract data for each.
I'll go through the OCR text sequentially.
After header, the text:
"Australia and
Avetr Ze land...}
British North
America,
British North
28
30,145 1,209|
:
:
}
7
8,578 248
:
28
30,145 1,209)||
7
8,578 248
14
16,531 207
132
17
Borneo,
1,065 95
:
3
1,065
95"
The "}" and "{" might be OCR artifacts for brackets. The numbers: 28, 30145, 1209, then 7, 8578, 248, then 28, 30145, 1209, then 7, 8578, 248, then 14, 16531, 207, then 132, 17.
This looks like two sets of British With Cargoes and In Ballast for two countries? But they are interleaved.
Perhaps the OCR read two columns side by side? The original might have had two pages or two columns. But the title says "No. 2" so it's one table.
Another idea: The "Australia and New Zealand" and "British North America" might be grouped under a region? But the numbers appear after "British North".
Let's assume the first country "Australia and New Zealand" has no data in the OCR (maybe zero), and the data starting with 28 belongs to "British North America". But then why is "Australia and New Zealand" listed?
Look at the totals: British With Cargoes vessels total 2,304. If British North America has 28, that's fine. Australia and New Zealand might have separate entry later? But not in list.
Maybe the OCR has a line: "Australia and New Zealand ... British North America" and the data for Australia is missing. The "28 30,145 1,209" might be for British North America. Then the next numbers "7 8,578 248" for its In Ballast. Then "28 30,145 1,209" repeated? Then "7 8,578 248" repeated? Then "14 16,531 207" for British Total? But 28+7=35 not 14. So not.
Perhaps the first group (28,30145,1209) is British With Cargoes for British North America. Second group (7,8578,248) is British In Ballast. Third group (28,30145,1209) is Foreign With Cargoes? But same numbers. Fourth group (7,8578,248) Foreign In Ballast? Fifth group (14,16531,207) Foreign Total? Then 132,17 for next country? But 132,17 is two numbers.
Then "Borneo" with "1,065 95 : 3 1,065 95". So for Borneo, British With Cargoes: 3 vessels, 1,065 tons, 95 crews. British In Ballast: none (colon). British Total: 3, 1,065, 95. Foreign With Cargoes: none.
But the "1,065 95" before colon might be a duplicate.
Let's move to "Coast of China and Formosa" which is clear:
"1,352 1,387,028 58,632 49 40,831 1,391 1,401 1,427,859 60,023 17,593 1,595,080 237,583"
Groups:
Good.
"Cochin China":
"48 34,971 1,348 63 53,439 1,905 111 88,410 3,253 53 36,023 1,334"
Groups:
"Continent of Europe":
"30 51,590 3,475 30 51,590 3,475 391 70,818 4,753"
Groups:
Only three groups. Perhaps British In Ballast is zero and not shown, so British Total = British With Cargoes (group1), but group2 is duplicate? Or group2 is British In Ballast (but same as WC). Group3 is Foreign With Cargoes. But then British Total missing. In totals, British Total vessels 2,484. For Continent of Europe, if British With Cargoes 30, and British In Ballast 0, British Total 30. But group2 is 30, so maybe group2 is British Total. Then group3 Foreign WC. So columns: Br WC, Br Tot, For WC. But other countries have Br IB.
"Great Britain":
"33 49,504 1,392 33 49,564 1,392 4,417) 131"
Groups:
Maybe group1 Br WC, group2 Br Tot (tons slightly different 49564 vs 49504), group3 For WC: 4417 vessels, 131? tons? crews? Missing one number.
"India & Singapore":
"139 176,480 7,278 14 13,298 326 153 189,778 7,004 31 36,304 1,228"
Groups:
"Japan":
"61 78,335 3,821 13 11,500 397 74 89,895 4,218 56 59,024 4,206"
Groups:
"Java and other Islands in the Indian Archipelago":
"1 510 19 3 2,849 3,359 96 8,504 670"
Groups:
Let's see: "1 510 19 3 2,849 3,359 96 8,504 670"
If groups of 3: (1,510,19), (3,2849,3359), (96,8504,670). That's three groups. Could be Br WC, Br Tot, For WC. Br IB missing.
"Macao":
"377 214,073 13,7309 176 25 379 214,249 13,764 974 1-45,258 24,652"
Parse numbers: 377, 214073, 137309? 13,7309 might be 137,309? Then 176, 25, 379, 214249, 13764, 974, 145258, 24652.
Groups: (377,214073,137309), (176,25,379), (214249,13764,974), (145258,24652) last incomplete.
But 176,25,379 - 25 tons for 176 vessels? Unlikely. Maybe 176 vessels, 25? no.
Perhaps the OCR has misread columns. "176 25 379" might be 176, 25, 379 for Br IB? But 25 tons too small.
"214,249 13,764 974" for Br Tot? 214,249 tons, 13,764 crews, 974? vessels? Order might be Vessels, Tons, Crews. So 214,249 vessels? No, tons. 13,764 crews, 974 vessels? Not consistent.
"1-45,258 24,652" might be 145,258 and 24,652 for For WC.
"Mauritius":
"1 300 38 1 1,300 38"
Groups: (1,300,38), (1,1300,38) - duplicate? Or Br WC and Br Tot.
"North Pacific":
"8 8,400 280 8. 8,490 280"
Groups: (8,8400,280), (8,8490,280)
"Pellew Islands":
"1 290 11 1 290 11"
Groups: (1,290,11), (1,290,11)
"Philippine Islands":
"26 12,470 900 16 11,388 269 41 23,858 1,169 33 12,899 1,083"
Groups:
"Ports in Hainan and the Gulf of Tonquin":
"91 34,097 2,566 3,660 100 96 37,757 2,720 32,975 1,987"
Groups: (91,34097,2566), (3660,100,96), (37757,2720), (32975,1987) - last two incomplete.
"Russia in Asia":
"1.200 173" - maybe 1,200 tons, 173 crews? Vessels missing. Or 1 vessel, 200 tons, 173 crews? "1.200" might be 1,200.
"Sandwich Islands":
"904 315" - two numbers.
"Siam":
"51 35,082 1,648 7 5,924 187 58 41,006 1,735 23 12,125 509"
Groups:
"South Africa": "A" - no data.
"South America": blank.
"South Pacific": blank.
"United States of America":
"56 93,496 3,451 498 181 +44 19 56 93,496 3,45 38.242"
Messy. Likely: 56, 93496, 3451 for Br WC. Then 498, 181, 44? for Br IB? Then 19, 56, 93496, 345? for Br Tot? Then 38,242 for For WC? Not clear.
"Totally":
"2,304 2,200.074 209,074 99,770 180 181,618 6,017 2,484 2,360,689 104.67 18,960 2,094,214 29,3T"
As parsed earlier.
Given the complexity, I will produce a markdown table with the 13 columns (Country, Br WC V, Br WC T, Br WC C, Br IB V, Br IB T, Br IB C, Br Tot V, Br Tot T, Br Tot C, For WC V, For WC T, For WC C). For each country, I'll fill in the numbers as they appear in the OCR in groups of 3, assuming the order: Br WC, Br IB, Br Tot, For WC. If a group is missing, leave blank. For countries with only 3 groups (9 numbers), I'll assume Br IB is missing (zero) and the groups are Br WC, Br Tot, For WC. But for consistency, I'll map based on the clear rows (Coast of China, Cochin China, India, Japan, Philippine, Siam) which have 4 groups (12 numbers) corresponding to Br WC, Br IB, Br Tot, For WC.
For rows with fewer numbers, I'll distribute as per pattern.
Let's define the columns:
| Countries | British With Cargoes Vessels | British With Cargoes Tons | British With Cargoes Crews | British In Ballast Vessels | British In Ballast Tons | British In Ballast Crews | British Total Vessels | British Total Tons | British Total Crews | Foreign With Cargoes Vessels | Foreign With Cargoes Tons | Foreign With Cargoes Crews |
|---|
Now fill row by row.
But the OCR has "28 30,145 1,209" right after "British North". So likely "British North America" is the first country with data. "Australia and New Zealand" might be a heading for a group? But it's listed as a country. In Hong Kong Blue Books, "Australia and New Zealand" is a country entry. I'll include it as a row with data missing (maybe zero). But the numbers 28,30145,1209 might belong to it? Let's see the text: "Australia and Avetr Ze land...} British North America, British North 28 30,145 1,209". The "}" might close a bracket for Australia. Then "British North America" is next. The "British North" might be a repeat. I'll assume the first data row is for "British North America". But then "Australia and New Zealand" is separate with no data. However, the totals row includes all. I'll list both.
Given the ambiguity, I'll follow the OCR sequence: first country mentioned is "Australia and New Zealand", then "British North America". The numbers start at "28". I'll assign the first set of numbers to "Australia and New Zealand"? But then "British North America" would have the next set. But the OCR shows "British North 28 30,145 1,209" - the "British North" might be part of the country name "British North America". So the country is "British North America" and the data follows. "Australia and New Zealand" might have been a country with no vessels cleared? But then why list it? In the original table, all countries are listed even if zero. The OCR might have lost the zeros.
I'll include both. For Australia and New Zealand, I'll put zeros or ...? The proofreading should preserve original content. If the original had zeros, they might not be printed. I'll leave blank or put 0.
But the user said: "Preserve original content — do not add or remove substantive words. Do not rephrase, rewrite, or change the author's style, tone, or word choice. Only correct what is clearly wrong." So I should not add zeros if not in OCR. I'll put ... for missing.
But the OCR has "Australia and New Zealand...}" and then numbers. The "}" might indicate a bracket for a group total. I'm not sure.
Given the time, I'll produce a table with the countries as they appear in the OCR in order, and for each, extract the numbers in groups of 3 as they appear in the text, assigning to the 12 data columns sequentially. Since the clear rows have 12 numbers for 4 categories, I'll assume each country has up to 12 numbers. For countries with fewer, the later columns are empty.
I'll write a script in my mind to parse the OCR text linearly. But manually.
Let me list the OCR text as a stream of tokens (numbers and country names).
Tokens:
Now, for each country, take the next 12 numbers (or as many as available) and assign to the 12 data columns in order: BrWC_V, BrWC_T, BrWC_C, BrIB_V, BrIB_T, BrIB_C, BrTot_V, BrTot_T, BrTot_C, ForWC_V, ForWC_T, ForWC_C.
But the numbers are not grouped per country in the stream; they are interleaved. The country names appear, then numbers. So for each country, the numbers following it until the next country name belong to it.
Let's segment by country names.
But "British North America" and "British North" might be the same. The numbers: 28, 30145, 1209, 7, 8578, 248, 28, 30145, 1209, 7, 8578, 248, 14, 16531, 207, 132, 17. That's 17 numbers. Too many. Next country is "Borneo". So British North America has 17 numbers? But we only have 12 columns. The extra 5 numbers (132,17) might be for Borneo? But Borneo has its own numbers.
Look: after 207, we have 132, 17, then "Borneo". So 132,17 might be the last two numbers for British North America? But 12 columns would take 12 numbers. Let's count: if we take first 12 numbers for British North America: 28,30145,1209, 7,8578,248, 28,30145,1209, 7,8578,248. That's 12 numbers. Then remaining 14,16531,207,132,17 are 5 numbers. But 14,16531,207 is a group of 3, and 132,17 is a group of 2. Then Borneo starts with 1065,95,3,1065,95.
This is messy. Perhaps the table has two sections: British and Foreign, and the OCR read columns in wrong order. But the user wants proofreading of the OCR output, not reconstruction of the original table from scratch. The instruction: "Return the proofread text in standard Markdown". So I should output the text with corrected OCR errors, preserving the structure as much as possible. The original text is a table. I should output a markdown table that represents the corrected data.
Given the difficulty, I'll produce a markdown table with the columns as per the header, and fill each row with the numbers
(34)
No. 2.-NUMBER, TONNAGE, and CREWS of VESSELS CLEARED at PORTS
COUNTRIES
TO WHICH DEPARTED.
WITH CARGOES.
Vessels. Tons.
BRITISH.
IN BALLAST.
TOT.L..
FOREIGN.
WITH CARGOES,
Crews. Vessels. Tons. Crews. Vessels. Tons. Crews. Vessels. Tons. Crews.
Australia and
Avetr Ze land...}
British North
America,
British North
28
30,145 1,209|
:
:
}
7
8,578 248
:
28
30,145 1,209)||
7
8,578 248
14
16,531 207
132
17
Borneo,
1,065 95
:
3
1,065
95
Coast of China
and Formosa...-
1,352 1,387,028 58,632
49 40,831 1,391 1,401 1,427,859 60,023 17,593 1,595,080|237,583|
Cochin China,......
48
34,971 1,348
63 53,439 1,905
111
88,410 3,253 53
36,023|| 1,334
Continent of
Europe,
30
51,590 3,475
30
51,590 3,475
391
70,818 4,753
:
Great Britain,............
33
49,504 1,392
33
49,564 1,392
4,417) 131
India & Singapore,
139
176,480 7,278 14 13,298 326
153
189,778 7,004
31
36,304 1,228
Japan,
61
78,335 3,821
13 11,500 397
74
89,895 4,218 56
59,024 4,206|
Java and other
Islands in the
1
510
Indian Archi-
19 3
2,849
3,359 96
8,504
670
pelago,
Macao,....
377
214,073 13,7309
to
176
25
379 214,249 13,764 974 1-45,258 24,652|
Mauritius,
1
300 38
:
1
:
1,300
38
:
North Pacific,
U
8 8,400 280
280
Pellew Islands,
1
290
11
:
1
290
11
:
:
:
÷
Philippine Islands, 26 12,470 900 16 11,388 269
41 23,858 1,169
33
12,899 1,083
Ports in Hainan
and the Gulf of Tonquin,.......
91 34,097 2,566
3,660 100
96
37,757 2,720
32,975 1,987
Russia in Asin,
:
***
:
***
:
1.200
173
Sandwich Islands,
:
..
904
315
Siam,.......
51 35,082) 1,648
7 5,924 187
58
41,006 1,735
23
12.125
509!
South Africa,
A
South America,
South Pacific,
United States of
America,
56
93,496 3,451
:
:
:
:
:
:
498
181
+44
:
19
56 93,496 3,45]||
38.242
Totally
2,304 2,200.074
209,074 99,770 180 181,618 6,017|2,484 2,360,689 104.67|| 18,9602,094,214 29,3T.
No comments yet.
Private notes are available after approval.