This is a guest piece by Aizat Ishembieva.
This post describes how structured text extraction and the creation of a page-referenced corpus can support research on nineteenth- and early twentieth-century travel writing about Central Asia.
My broader research examines how Central Asian peoples, places, livelihoods, and material culture were observed and described through Western lenses during the late nineteenth and early twentieth centuries. It focuses particularly on the Danish Pamir Expeditions (1896–1897; 1898–1899), and how the expedition members’ encounters with Central Asian peoples, places, and material culture were transformed into travel narratives, ethnographic observations, collected objects, and scholarly knowledge. It also traces how these historically situated classifications were later preserved, interpreted, or reworked in museum catalogues and digital databases.
In two posts published last year, I explored this process through the Olufsen Collection. Part 1 traced how objects collected during the Danish Pamir Expeditions were registered and later represented in the National Museum of Denmark’s online database. Part 2 compared that database with Esther Fihl’s Exploring Central Asia, showing how a digital database and a scholarly catalogue provide different forms of access and interpretation. My present pilot research has shifted from the material collection and its metadata, to Ole Olufsen’s published works. His use of terminology and descriptive choices provide evidence of how he presented Central Asia to readers. English-language accounts by Ármin Vámbéry, Eugene Schuyler, and Francis Younghusband provide comparative sources for identifying shared patterns and differences between authors.
To support my new angle of enquiry, I have used structured text extraction to build a page-referenced corpus and used this as the source for an experimental research assistant in Microsoft Foundry. The questions that guide the following articles and my reflection on this software are as follows:
- How effectively can a Microsoft Foundry research assistant help historians locate relevant passages, compare descriptive language across travel accounts, and trace evidence back to the corresponding pages in the digitised source books?
- What limitations arise from OCR errors, missing pagination, retrieval failures, and inaccurate model-generated answers?
Azure RAG implementation in Microsoft Foundry
My project implemented an experimental, corpus-grounded research assistant in Microsoft Foundry. The system comprised the Aizat-Books-RAG-Assistant agent (version 7), a Foundry IQ knowledge base supported by Azure AI Search, managed-identity connections, role-based access control, and GPT-4o-mini for answer generation.
The comparative corpus contained five published works in six volumes:
- Ole Olufsen, Through the Unknown Pamirs (1904);
- Ole Olufsen, The Emir of Bokhara and His Country (1911);
- Ármin Vámbéry, Travels in Central Asia (1864);
- Eugene Schuyler, Turkistan: Notes of a Journey in Russian Turkistan, Khokand, Bukhara, and Kuldja, Volumes I and II (1877 edition; copyright 1876);
- Francis E. Younghusband, The Heart of a Continent: A Narrative of Travels in Manchuria, across the Gobi Desert, through the Himalayas, the Pamirs, and Chitral, 1884–1894 (1896).
Preparing the Page-Aware JSON Corpus
Page-aware JSON files were first uploaded to Microsoft Foundry – “page-aware” is a project-specific term rather than a standard file format. Each JSON record represented one PDF page and combined the extracted text with the author, title, publication year, PDF page number, printed page number where identifiable, source identifier, record ID, and a prepared citation. The unique record ID was intended to connect each retrieved passage to a specific page and book.
PDF and printed pagination were recorded separately. The PDF page identifies the page’s position in the digital file, while the printed page is the number visible in the original book. These numbers can differ because digitised volumes may contain covers, front matter, maps, illustrations, inserted plates, and unnumbered pages. When a printed page number could not be verified, it was left unspecified rather than inferred.

Figure 1. Example of a page-aware JSON record containing extracted text and source metadata.
Ingestion and retrieval
The page-aware JSON files were then uploaded as source material for a Foundry IQ knowledge base. A knowledge base is not the dataset itself, it is an orchestration layer that connects an agent to one or more knowledge sources while Azure AI Search provides the underlying indexing and retrieval infrastructure (Microsoft 2026d).
The simplified workflow I followed was:
User query → Foundry agent → Foundry IQ knowledge base → Azure AI Search → retrieved passages → gpt-4o-mini → generated answer
Figure 2. Simplified workflow of the experimental Microsoft Foundry implementation (picture generated by AI).

Figure 3. Microsoft Foundry. Aizat-Books-RAG-Assistant.
This representation does however omit some configuration-dependent operations. Agentic retrieval can divide a complex request into subqueries, execute keyword, vector, or hybrid searches, rerank the results, and combine the retrieved evidence (Microsoft 2026a).
The assistant was instructed to treat the historical corpus as its primary source of evidence, to avoid unsupported claims and invented citations, identify the source of retrieved information, and state when sufficient evidence could not be found. Source-specific questions were restricted to the named work, while comparative questions were divided into separate searches for each author or book. The instructions also included historical spelling variants and possible OCR variations, such as Kirghiz/Kyrgyz/Kirghis, Bokhara/Bukhara, and Tadjik/Tajik, etc.
Only the page-aware JSON files were uploaded for this implementation. The files entered Azure’s managed ingestion and indexing workflow. Microsoft documents that Foundry IQ can automate document chunking, vector-embedding generation, and metadata extraction. However, this general capability does not establish exactly how the JSON files were processed in this particular implementation (Microsoft 2026d).
Results depended on how Azure parsed and mapped the JSON files, but the ingestion settings used in this project were not verified (Microsoft 2025). Similarly, the presence of text-embedding-3-small does not prove that vector retrieval was active; this requires the model to be properly connected to the search index (Microsoft 2026b).
But still, further configuration checks are needed to confirm how Azure parsed and indexed the JSON files, whether page boundaries and metadata were preserved, which retrieval method was used, and whether citations pointed to the correct source records. Page references returned by the chatbot should therefore be checked against both the uploaded page-aware JSON records and the original PDFs before being used as scholarly citations.
The resulting chatbot should be presented as an experimental RAG implementation based on page-aware JSON files. This distinction is also important because, at the time of writing, access through the Microsoft Foundry and Azure portals to all agentic-retrieval features remained preview-only (Microsoft 2026c).
Preliminary evaluation example
Of course, to test a functioning chatbot as a reliable historical research instrument, it must undergo preliminary manual evaluation. Here is an example of a research question being tested by comparing the chatbot’s answer with a human-verified passage from the source.
The test query was:
“How does Olufsen describe the livelihoods and settlement patterns of the people he referred to as “mountain Tadjiks” and “Kirghiz” in the Shakhdarra Valley? ”
And the chatbot’s generated response:

Transcript: Mountain Tadjiks: They primarily engage in agriculture, cultivating fields and gardens on terraces formed from the valley’s loose sediment. They use artificial irrigation to till the soil, particularly at lower altitudes, where they can grow various crops like wheat and melons12. Their villages, or “kishlaks,” are situated along riverbanks and on terraces, producing a relatively stable agricultural lifestyle3. While they sometimes migrate with their cattle to higher pastures during summer, their way of life remains largely settled and agrarian45.
Kirghiz: In contrast, the Kirghiz are predominantly nomadic pastoralists. Their activities center around the breeding of cattle and sheep, with a focus on finding fresh pastures6. They migrate seasonally between designated winter and summer pastures, living in yurts and portable structures, which contrasts sharply with the Tadjiks’ more permanent settlements7. Although some Kirghiz families may settle temporarily in the Shakhdarra valley during summer, their lifestyle remains inherently nomadic8.
These differences highlight a clear distinction in how these two groups interact with their environment, with the Tadjiks favoring a settled, agricultural lifestyle and the Kirghiz maintaining a nomadic existence based on livestock herding910.
The chatbot responded that Olufsen distinguishes the two groups through their agricultural and pastoral practices. It described the mountain Tadjiks as primarily settled agriculturalists living in kishlaks and using artificial irrigation, while also moving their cattle to higher pastures during the summer. The Kirghiz were described as nomadic pastoralists who moved seasonally between pastures and used the Shakhdarra river valley temporarily during the summer.
The response identified the main contrast reasonably well. It correctly stated that the mountain Tadjiks lived in permanent settlements, practised irrigated agriculture, and also moved their livestock seasonally. It also recognised that some Kirghiz (Kyrgyz) families used the higher territory as a summer nomadic area. However, the answer included several details that are not found in the human-verified Shakhdarra passage, including references to terraces formed from loose sediment, the cultivation of wheat and melons, villages situated along riverbanks, yurts, sheep breeding and movement between designated winter and summer pastures.
The chatbot also omitted important evidence from the verified passage. It did not mention that Seis, at approximately 3,300 metres, marks the upper boundary of the agricultural area; that cultivation ends above this point; that the stone huts in the higher area are interpreted as seasonal Ailäks rather than ruined permanent settlements; or that Olufsen uses Iranian and Turkish or Kirghiz place-names as evidence when distinguishing the historical use of the territory. The most serious problem is the broad ethnic generalisation. Olufsen states that the area near the Mas Pass was used by “some Kirghiz families,” whereas the chatbot describes the Kirghiz generally as having an “inherently nomadic” lifestyle.
Conclusion
This single preliminary example suggests that the Microsoft Foundry can broadly address the relevant subject and produce a partly accurate synthesis. However, the generated answer may omit important details and broaden specific observations. The evaluation is still in progress, and no general conclusions can yet be drawn about the chatbot’s overall reliability. Historical OCR, spelling and transliteration remain challenging. Automatic modernisation may improve retrieval but would remove historically significant evidence, so uncertain passages must be checked against the original scans. Overall, Microsoft Foundry does offer an interesting avenue for this type of textual extraction and analysis within its current limitations.
References:
Ole Olufsen, Through the Unknown Pamirs: The Second Danish Pamir Expedition, 1898–99 (London: William Heinemann, 1904).
Ole Olufsen, The Emir of Bokhara and His Country: Journeys and Studies in Bokhara, with a Chapter on My Voyage on the Amu Darya to Khiva (Copenhagen: Gyldendalske Boghandel, Nordisk Forlag; London: William Heinemann, 1911).
Eugene Schuyler, Turkistan: Notes of a Journey in Russian Turkistan, Khokand, Bukhara, and Kuldja. 2 vols (New York: Scribner, Armstrong & Co., 1877).
Ármin Vámbéry, Travels in Central Asia (London: John Murray, 1864).
Francis E. Younghusband, The Heart of a Continent: A Narrative of Travels in Manchuria, across the Gobi Desert, through the Himalayas, the Pamirs, and Chitral, 1884–1894 (London: John Murray, 1896).
Microsoft, “Index JSON Blobs and Files in Azure AI Search.” Microsoft Learn (2025). Last updated October 9, 2025. https://learn.microsoft.com/en-us/azure/search/search-how-to-index-azure-blob-json.
Microsoft, “Agentic Retrieval in Azure AI Search.” Microsoft Learn (2026a) Last updated June 12, 2026. https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview.
Microsoft, “Azure OpenAI Vectorizer.” Microsoft Learn (2026b). Last updated July 30, 2026. https://learn.microsoft.com/en-us/azure/search/vector-search-vectorizer-azure-open-ai.
Microsoft, “Create a Knowledge Base in Azure AI Search.” Microsoft Learn (2026c). Last updated August 31, 2026. https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-how-to-create-knowledge-base.
Microsoft, “What Is Foundry IQ?” Microsoft Learn (2026d). Last updated August 1, 2026. https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq.
AI Disclosure: Microsoft Foundry, using GPT-4o mini, was used to generate answers to questions about the historical travel book corpus. The answer discussed in the article was manually checked against the source passage, identifying unsupported details and omissions. Moreover, ChatGPT (GPT-5.6) was used to assist with writing and troubleshooting Python extraction scripts, and create the workflow illustration (Figure 2). The author ran and tested the scripts locally in Visual Studio Code, using Docling 2.115.0 and EasyOCR 1.7.2., and reviewed the illustration.

