Ohlone Language Learning
Ohlone Language Atlas

A Chochenyo-first learning tool with source-backed words, pronunciation, and practice.

Corpus Workbench

Grow the Ohlone language corpus with reviewable source and audio queues.

The dictionary is useful now, but the next jump is provenance: item-level sources, approved text examples, and phrase-level audio alignment that can support a cautious AI assistant.

Current Corpus
The data is broad enough for retrieval, not for unsupervised generation.
Use the corpus to cite, compare, and explain attested forms. Training speech or story generation needs permissioned aligned data.
1959 entries42 phrases16 sources28 audio leads
Coverage

Entries by variety

8 tracked varieties
Awaswas

390

dictionary entries

Chochenyo

112

dictionary entries

Mutsun

413

dictionary entries

OCEN Rumsen

47

dictionary entries

Ramaytush

112

dictionary entries

Rumsen

593

dictionary entries

Santa Clara

140

dictionary entries

Soledad

152

dictionary entries

Written Source Queue
Prioritize sources with text, translations, and item-level references.
10 high-priority written sources are waiting for review or deeper itemization.
high priorityscraped full text

Mutsun-English English-Mutsun Dictionary PDF

This is a 677-page text-readable dictionary with CC BY-NC 4.0 licensing and enough structure to become the primary Mutsun lexical source.

Next: Parse the scraped page/chunk files into reviewed dictionary entries and examples.

high prioritypartially represented

Mutsun Text Collection

Longer texts are more valuable than isolated word lists for grammar, examples, and RAG context.

Next: Inventory each downloadable text and map it to phrase/example records with page or item references.

medium prioritymetadata only blocked by 403

Quizlet: Chochenyo final set (all words)

Search indexing reports a 188-card Chochenyo learning set with greetings, pronouns, and everyday vocabulary, which could help identify modern learner forms to verify against primary/community sources.

Next: Do not bulk scrape. Ask the set author or relevant Chochenyo/Muwekma language authority for permission, then verify every card against primary or community-approved sources before ingestion.

high prioritynot ingested

Mutsun Language Database Exports

Potentially structured lexical and text material that may reduce manual cleanup.

Next: Check download access, license, and field schema before import.

high prioritynot ingested

Pinart mission vocabularies / Heizer 1952

Digitized mission-era Costanoan wordlists may fill gaps for Chalon, Awaswas, Rumsen, and related historical varieties.

Next: Review edition rights and community sensitivity, then extract item-level vocabulary with historical-transcription status.

high prioritynot ingested

Schoolcraft Costanos vocabulary

One of the few online lexical leads for the Mission Dolores / San Francisco Bay area vocabulary.

Next: Extract only after Ramaytush review; label forms as early colonial transcription.

Audio Alignment Queue
The scarce resource is clean audio-text alignment.
15 high-priority media leads should be checked for captions, written forms, and permissions first.
high priorityChochenyonot started

Chochenyo Welcome

Captions: auto captions only observed

Next: Check if the source exposes written Chochenyo text or a reusable transcript.

medium priorityChochenyonot started

Jose Guzman Muwekma Chochenyo Song

Captions: none observed

Next: Use as catalog reference only until community reuse permission is explicit.

high priorityRumsennot started

Reviving the Ohlone Language

Captions: manual english observed

Next: Ask Smithsonian and speaker/community before using; manually transcribe only approved Ohlone spans.

Build Path

Make RAG useful first, then decide if training is justified.

Step 1

Inventory sources

Step 2

Review rights

Step 3

Extract text

Step 4

Align audio

Step 5

Ground assistant

The website should remain language-centered. History, maps, songs, stories, and archival context should enter through source-backed language records instead of becoming a broad general-history site.