How Ask the Archives Works

Getting Started

Ask the Archives is an AI-powered research assistant that helps you explore historical documents in our collections. Simply type a question in natural language, and the system will search through transcribed documents to find relevant information and provide you with a summary.

You can ask follow-up questions, and the AI will remember the context of your conversation to provide more relevant responses.

How Your Question Is Processed

  1. Keyword Analysis: When you ask a question, the AI analyzes your query to identify key terms, names, dates, and concepts that might appear in historical documents.
  2. Database Search: The system searches a large database of document transcriptions using these keywords and related terms to find potentially relevant material.
  3. Content Analysis: The AI reads through the matching transcriptions to understand their content and relevance to your question.
  4. Response Generation: Finally, the AI synthesizes what it found and presents a summary, with citations linking to the original source documents.

Why can't the AI just read everything? Our collections contain thousands of pages of historical documents. Because of this large quantity of content, the AI cannot process and hold everything in memory at once. Instead, it must strategically search for relevant material based on your specific question.

Understanding Citations

When the AI references specific documents, you'll see numbered citations in the response. Hover over a citation to see details about the source, including the document title and page. Click the link to view the original document image.

Tips for Better Results

  • Be specific: Include names, dates, places, or events when you know them.
  • Try different phrasings: Historical documents may use different terminology than modern language.
  • Ask follow-up questions: If the first response mentions something interesting, ask for more details.
  • Start broad, then narrow: Begin with a general question, then focus on specific aspects in follow-ups.

Important Limitations

Please keep these factors in mind when using Ask the Archives:

  • Not everything is transcribed: Only documents that have been transcribed are searchable. Many items in our collections have not yet been processed.
  • Transcription accuracy varies: Historical documents can be difficult to read due to handwriting, age, or damage. Transcriptions may contain errors.
  • AI interpretation: The AI does its best to understand and explain historical content, but it may occasionally misinterpret unfamiliar terms or context.

For Researchers: There are multiple levels of verification needed to validate information found through this tool. Always consult the original document images when conducting serious research. The AI is excellent at summarizing content and providing a starting point, but should not be treated as a definitive source.

Technical Details

Ask the Archives uses a retrieval-augmented generation (RAG) approach. Your question is used to query a vector database of document transcriptions. The most relevant passages are retrieved and provided to a large language model along with your question and conversation history. The model then generates a response based on this context.

Citations use a special markup format that links responses to specific document images in our digital collections. This allows you to verify information against the original source material.

Conversation history is maintained in your browser session, allowing for contextual follow-up questions. Starting a "New Chat" clears this history and begins a fresh conversation.

AI System Instructions

The AI is given these instructions when responding to your questions:

  • You are an expert historian helping interested lay people understand the history told in these documents. The documents are almost entirely records from early religious congregations in Philadelphia, mostly Christian but also including one Jewish congregation (Mikveh Israel).
  • CITATION REQUIREMENTS - You MUST follow these rules for EVERY response:
  • 1. EVERY factual claim about the archives MUST include a primary source citation in the format [[image_id:12345]]. Do not make claims about archive contents without citing the specific image.
  • 2. EVERY response MUST include web citations for historical context using the format [[web:Source Title|URL]]. Example: [[web:William Penn - Wikipedia|https://en.wikipedia.org/wiki/William_Penn]]
  • 3. A complete response will have BOTH types of citations: primary sources ([[image_id:...]]) for archive-specific claims AND web sources ([[web:...|...]]) for broader historical context.
  • PRIMARY SOURCES: Base all specific claims about archive contents on the primary sources. Include short quotes when they add clarity or color. If you cannot find support in the primary sources for a claim, say so explicitly.
  • WEB SEARCH: Always search for broader historical context to enrich your answers. Use Wikipedia, the Encyclopedia of Greater Philadelphia, congregation websites (Christ Church, St. Peter's, Gloria Dei, Mikveh Israel), ushistory.org, and National Park Service resources. Web sources provide the historical backdrop that helps readers understand the primary sources.
  • RESPONSE STRUCTURE: Lead with insights from the primary sources, then weave in historical context from web sources. Every paragraph making factual claims should include at least one citation.

Questions or Feedback?

If you encounter issues or have suggestions for improving this tool, please contact the archives staff. We're continually working to expand our transcribed content and improve the search experience.