WikiJournal Preprints/OpenSpeaks: Open Toolkit for Multimedia Documentation of Indigenous Languages
This article is an unpublished pre-print undergoing public peer review organised by the WikiJournal of Humanities.
You can follow its progress through the peer review process at this tracking page.First submitted:
QID: Q106806074
Suggested (provisional) preprint citation format:
Subhashish Panigrahi. "OpenSpeaks: Open Toolkit for Multimedia Documentation of Indigenous Languages". WikiJournal Preprints. Wikidata Q106806074.
License:
This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction, provided the original author and source are credited.
Eystein Thanisch
contact
Sim Tze Wei
T. B. Dinesh
Article information
Abstract
Indigenous, endangered, and other low‑resource languages continue to be documented and archived through practices that often reproduce colonial control over data, rights, and access for native speakers. Citizen‑led language documentation, especially by non‑native speakers, must instead centre community agency as a core design principle. This article presents OpenSpeaks, an open educational toolkit as a combined framework for planning, recording, describing, and publishing multimedia language documentation with and for communities. It draws on citizen‑led documentation and archiving experiences to outline a practical workflow that foregrounds consent, content rights, open licensing, and accessibility, and that treats archives as shared, long‑term community infrastructure rather than external repositories. The toolkit, hosted on Wikiversity, and the community‑driven OpenSpeaks Archives support activists, community documenters, and archivists working with low‑resource languages, particularly those that lack prior audiovisual documentation. From an anthropological and social‑justice perspective, the article reflects on gaps and tensions—around informed consent, collective ownership, disability access, and speech‑technology readiness—and proposes a citizen‑driven, open, and accessible model for multimedia documentation.
Introduction
[edit | edit source]The rapid decline of Indigenous, endangered, and other low‑resourced languages is widely recognised as a loss of human knowledge and cultural diversity.[1] Many of these languages receive little or no sustained documentation, and the few existing records often remain inaccessible to the communities whose voices they contain. Activists and citizen archivists have therefore turned to digital media and networked platforms to document everyday speech, stories, and songs, and to share them under open licenses that allow reuse in education, revitalisation, and research.[2]
Digital interventions, such as online community media, subtitled videos, and social‑media‑based campaigns, can normalise the public use of Indigenous and minoritised languages and create a ripple effect into offline life.[3] The spread of affordable smartphones and low‑cost recording equipment has made audiovisual documentation more feasible for local initiatives, though connectivity, hardware, and skills gaps remain unevenly distributed.[4] At the same time, documentary linguistics continues to stress the value of naturalistic recordings of face‑to‑face interaction for long‑term archives and future research.[4]
Initiatives like Rising Voices show that younger speakers and community activists can lead digital language revitalisation when they receive targeted support, clear workflows, and reusable resources.[5] However, new documenters often struggle to navigate questions of consent, content rights, licensing, accessibility, metadata, and long‑term preservation.
OpenSpeaks was created in 2017 as a standalone project and later moved to Wikiversity as an open educational resource (OER) for citizen language archivists.[6][7] In 2025, OpenSpeaks Archives was launched in 2024 as an open and public multimedia archive optimised for Wikipedia and Wikimedia projects, starting with descriptive speech recordings and oral histories in nearly 20 South Asian languages.[8] This article uses “OpenSpeaks” to refer to the toolkit and “OpenSpeaks Archives” for the archive, and introduces the combined OpenSpeaks Framework that now guides both.
Background
[edit | edit source]Language documentation and archiving are no longer limited to field linguists; they involve activists, community content creators, educators, and archivists.[4] Community‑driven documentation can create new media that both preserve knowledge and support present‑day language use, but it also inherits longstanding inequalities around who controls archives and who can access recordings.[9]
In many contexts, textual documentation is constrained by the lack of widely used writing systems, missing or partial Unicode support, rendering issues, and poor availability of fonts, devices, and literacy in the language.[10] Audiovisual recording can therefore be a more appropriate starting point for documenting highly oral languages, provided that local constraints around equipment, space, safety, and time are carefully considered.[4]
Different documentation practices have different priorities: a documentary linguist might seek dense linguistic data, a filmmaker might emphasise narrative and aesthetics, and a health campaign might treat language documentation as a by‑product of awareness videos.[11] OpenSpeaks deliberately takes a cross‑disciplinary view, drawing from documentary linguistics, community media, and digital rights work to make practical guidance that can be adapted by different stakeholder groups.
Oral History Framework
[edit | edit source]Earlier drafts of this work used the term “Citizen Language Documentation and Archiving (CLDA)” to describe community‑based audiovisual documentation inspired by citizen science.[12] Experience with OpenSpeaks and OpenSpeaks Archives showed that this framing over‑emphasised the external archivist and did not distinguish clearly between different kinds of documenters.
The current OpenSpeaks Framework explicitly recognises three overlapping roles:
- Community documenters: community members recording their own language, culture, and histories, often for intra‑community use and intergenerational transmission.
- Neighbouring Indigenous documenters: documenters from nearby or related Indigenous communities who collaborate as peers.
- External non‑Indigenous documenters: linguists, researchers, journalists, or allies who document and support archiving from outside the community.
Across these roles, OpenSpeaks emphasises “Nothing About Us Without Us” as a practical principle that shifts decision‑making on consent, content rights, licensing, and access to the people whose voices are recorded.[13] The Framework also assumes that language documentation is iterative: recordings, consent decisions, and metadata can evolve over time as communities revisit agreements and repurpose their archives.
OpenSpeaks Toolkit and Archives
[edit | edit source]Toolkit: chapters and scope
[edit | edit source]On Wikiversity, OpenSpeaks is currently organised into four interrelated chapters designed for beginner and intermediate archivists: (a) consent, content rights and open licensing, (b) audiovisual recording, (c) metadata collection and publication, and (d) accessibility.[7] The text uses plain English, short sections, and editable templates so that communities can translate and adapt the material to their own contexts,[6] such as the first chapter which was translated into Santali.[14]
Version 1.0 focused mainly on practical recording steps and basic pre‑production and post‑production, while Version 2.0 expanded consent, content rights, and licensing guidance and introduced a reusable content release template.[15] Version 3.0 and later work, supported in part by Wikimedia Deutschland’s UNLOCK programme, integrates accessibility and disability‑centred design throughout the workflow, including guidance on subtitles, captions, transcripts, and screen‑reader‑friendly descriptions.[5]
OpenSpeaks Archives: goals and pilot
[edit | edit source]OpenSpeaks Archives extends the toolkit into a working archive that prioritises descriptive recordings of natural speech, embeddable and citable across Wikimedia projects and beyond.[16] The pilot focused on five tongues, namely, Kusunda, Baleswari Odia, Bonda, Ho, and Van Gujjari, and produced professionally edited videos with:
- descriptive, unscripted speech on everyday topics
- clean audio (no background music except when documenting music)
- subtitles in at least one official/local language and English
- closed captions and transcripts where possible
- structured metadata and links to Wikidata items, Commons files, and relevant Wikipedia articles.[17]
These recordings have been integrated into Wikipedia, Wiktionary, Wikisource, Wikidata, and Wikimedia Commons in more than twenty language editions.[8]
The Archives workflow builds on GLAM (Galleries, Libraries, Archives, Museums) practices:
- oral history recordings are acquired by GLAM institutions and catalogued, and are citable on Wikimedia projects as a form of Indigenous knowledge
- metadata is aligned with Wikidata and other open databases
- content is curated for long‑term preservation through content negotiation with partner institutions.[8]
Consent
[edit | edit source]OpenSpeaks treats consent as an ongoing, relational process rather than a one‑time form, especially in low‑literacy or highly unequal settings.[18] The toolkit encourages archivists to explain in clear, local language:
- what will be recorded
- how the recording might be edited and shared
- whether it will be uploaded to open platforms such as Wikimedia Commons
- what licenses may apply and what they imply for future reuse.
To reflect peer feedback, OpenSpeaks now stresses that written consent alone is insufficient when participants have limited literacy in the language of the form or limited experience with digital media. Instead, consent conversations are treated as iterative, with room for questions, clarifications, and community‑level decision‑making about sensitive uses.
The OpenSpeaks content release template offers a simple, adaptable text that combines consent and rights release while allowing projects to specify the exact Creative Commons license or other arrangement that works for the community.[19] In practice, many interviews record verbal consent on video or audio, and then summarise key points in metadata and accompanying documentation.
Content rights, copyright, and licensing
[edit | edit source]Ownership of recorded content can be complex, especially when individual narrators, communities, and commissioning organisations all have legitimate interests in how stories and knowledge are shared.[20] OpenSpeaks uses case‑based prompts rather than fixed rules, asking archivists to identify:
- whether a recording captures collective knowledge (e.g. a well‑known song or story) or an individual creation
- whether the work is self‑initiated, commissioned, or part of a partnership
- which agreements, if any, already exist about copyright and reuse.
Accessibility and annotations
[edit | edit source]Accessibility requirements need to provide disability access as well as the accessibility of content and tools for the documented communities themselves. Hence, OpenSpeaks integrates accessibility at several levels:
- Technical: ensuring audio clarity, avoiding background music where it would interfere with speech recognition, and using open standards and formats (e.g. WebM, WAVE, FLAC).
- Linguistic: providing subtitles and/or transcripts in the recorded language where possible, and/or at least one widely used language for broader access.
- Descriptive: encouraging meaningful annotations that explain culturally specific practices, objects, or events shown in the recording, where possible.
One yet-to-be implemented consideration, if resources permit, is whether annotations themselves will require their own consent and privacy considerations when they add sensitive contextual information. The annotation workflows will need to include
- clear attribution of annotators in metadata
- explicit discussion of what kinds of annotations should be public, restricted, or anonymised
- storage of annotations in both personal and public archives with appropriate safeguards
Audiovisual recording: updated guidance
[edit | edit source]Chapter 2 of OpenSpeaks provides a step‑by‑step overview of planning and conducting audiovisual recordings in field and home‑studio conditions.
Prerequisites and environment
[edit | edit source]Key preparation steps now include:
- agreeing in advance on the quietest possible location and noting unavoidable background sounds in metadata
- checking that any video or image without personally identifiable information (e.g. human faces, names, addresses, copyrighted image/sound or precise locations) is treated differently from recordings that do contain such information
- where recordings with identifiable information are essential, ensuring that all personally identifiable information that is not necessary is redacted or masked in the final edit
The recording environment section retains practical tips (using lavalier microphones, avoiding flickering LED lighting, preferring natural light) and explicitly notes that many modern cameras and smartphones include built‑in stabilisation that can complement low‑cost tripods.
Audio and video quality
[edit | edit source]The audio section continues to prioritise clear, lossless audio, encouraging recording at 44.1 kHz or higher, with at least 16‑bit depth, and saving in lossless formats such as WAV or FLAC where feasible. The guidance now clarifies missing wording and stresses that bundled earphones usually offer better audio than device‑internal microphones when used in quiet environments.
For video, OpenSpeaks recommends recording at a minimum of Full high-definition (HD) i.e. a resolution of 1920×1080 pixels, where devices allow, so that viewers and researchers can observe mouth movements and subtle gestures that are part of communication. The guide stresses that:
- audio quality remains more critical than very high‑resolution video for most documentation scenarios
- when budgets are limited, investing in a reliable microphone is often more impactful than upgrading cameras
- device default microphones are rarely sufficient for long‑term documentation
Metadata, subtitles, and publication
[edit | edit source]Chapter 3 provides downloadable templates for metadata sheets and content release forms and now links more explicitly to Wikimedia Commons upload workflows and structured data practices.[19] It highlights that:
- subtitles and captions should be time‑aligned, using open formats such as WebVTT or SRT or TimedText
- transcriptions should be prepared with native speakers wherever possible
- summaries can make long transcriptions more approachable for community members who may not wish to read full verbatim texts
The chapter also notes that, when preparing media for Wikimedia projects, archivists should ensure that filenames, descriptions, and category structures make recordings discoverable as language resources, not only as audio or video files.
Learning exercises and localisation
[edit | edit source]During the development of Version 2.0, a bilingual (English–Santali) survey with 25 Santali‑speaking practitioners offered a first systematic check of how emerging community documenters understand consent, content rights, and licensing in practice.[15] The sample was small but diverse: 17 respondents were native Santali speakers spread across five countries; around one‑third represented collectives, nonprofits, academic or other civil society organisations, while the rest practised digital activism in a personal capacity; and almost half were also archiving at least one additional Indigenous language, including several with no formally recognised writing system.[15] To summarise, many community documenters already operate trans‑locally, across multiple languages and platforms, but still lack targeted guidance on ethics and rights.
On consent, over half of the respondents reported not being fully confident about how to ask for consent during documentation, even though most had already been recording and publishing content.[15] Participants described three main consent practices: informal verbal discussions before recording, consent negotiated “on camera” during the recording, and written forms filled in before an interview. Most reported relying on verbal or in‑recording consent rather than written forms, which aligns with contexts where literacy in the language of the form and familiarity with legal documents are uneven. For the toolkit, this confirmed that OpenSpeaks needed to treat consent as an ongoing conversation, not a checkbox, and to explicitly legitimise verbal and audio‑recorded consent where communities prefer these forms, while still documenting decisions clearly in metadata.
On copyright and licensing, almost half of the respondents said they knew how to make audiovisual recordings in their languages but needed a beginner‑friendly guide to “what happens next” with rights and reuse.[15] Most highlighted a lack of practical understanding of copyright basics and Creative Commons licenses, even when they were already uploading content to social media or Wikimedia projects. For OpenSpeaks, this validated the decision to devote a full chapter to consent, content rights, and licensing, structured around concrete scenarios, questions, and trade‑offs rather than abstract legal explanation. It also reinforced the choice, later adopted in OpenSpeaks Archives, to keep license options simple, clearly explain which licenses are Wikimedia‑compatible, and foreground implications for community access and possible revocation.
The localisation process for Santali showed that translation could not be a word‑for‑word exercise. The Santali translation team included Wikimedians, R. Ashwani Banjan Murmu, Fagu Baskey, and Joy Sagar Murmu, who combined loanwords (for instance, transliterations of terms like “license” into the Ol Chiki script), newly coined expressions, and existing vocabulary to explain concepts such as open licensing, attribution, and public domain in ways that felt natural to local readers.[21] Their feedback loop into the English chapter helped break the earlier, linear, “legal‑first” structure of Version 2.0: the revised text starts from common dilemmas that survey participants raised and then introduces terms and tools as needed. This bidirectional influence—survey to translation to English base text—is now a template for how OpenSpeaks aims to localise and co‑design materials with other language communities.
Subsequent workshops at WikiConference India, Celtic Knot, and other Wikimedia events extended these exercises into practical sessions where participants used OpenSpeaks resources to plan and record short field interviews, reflecting on consent, rights, and metadata as they worked.[22] These activities continue to refine the OpenSpeaks Framework, informing the design of OpenSpeaks Archives and collaborative campaigns such as Wiki Loves Languages.
Implications for documentary linguistics and language technology
[edit | edit source]The Oral History Framework and OpenSpeaks Archives position community‑led, ethically grounded audiovisual documentation as a foundation for both language justice and future technologies. For speech technologies in particular, community‑controlled descriptive recordings offer more diverse training data than small, curated studio corpora alone, provided that data collection is fair, consentful, and inclusive across gender, age, and socio‑economic groups, and most importantly, ensuring of community control through consent and license agreements.[23][13]
For documentary linguistics, OpenSpeaks adds a practice‑oriented, community‑first perspective to existing technical and methodological guides, showing how simple tools and workflows can align with high archival standards when consent, collective ownership and rights, and accessibility are treated as primary design principles rather than afterthoughts.[4][24] By publishing its toolkit as an OER and aligning its archive with Wikimedia infrastructures, OpenSpeaks also invites communities, educators, and researchers to fork, adapt, and extend its materials to their own circumstances.
Additional information
[edit | edit source]Self declaration
[edit | edit source]The author self‑identifies as a dominant‑caste, cisgender male individual and acknowledges that this socio-economic location has shaped his access to education, English, digital technologies, and international networks. These privileges have influenced his early participation in open knowledge and Wikimedia movements and his ability to design and maintain OpenSpeaks and OpenSpeaks Archives. He recognises that such positionality carries risks of reproducing colonial, caste‑based, and patriarchal structures even within projects that aim to support Indigenous and marginalised communities.
Acknowledgements
[edit | edit source]OpenSpeaks has been enriched from a range of major projects, readings, and interactions. It might not be possible to attribute all in a chronological order but some of the individuals and organizations include, but is not limited to:
- Indigenous communities: Santali community (specifically Ramjit Tudu, R Ashwani Banjan Murmu, Fagu Baskey and Joy sagar Murmu); Bonda community of Bandhuguda, Malkangiri district, Odisha, India; Ho community of Keshpada, Mayurbhanj district, Odisha, India; Kusunda, Tharu and Magar communities of Kulmor, Dang district, Nepal; Gutob community of Tukum, Koraput district, Odisha, India
- Civil society partners/donors: Eddie Avila and Rising VoicesGlobal Voices, Creative Commons, National Geographic Society, Mozilla Open Leadership Series, MJ Bear Fellowship 2017, Online News Association, WhoseKnowledge?, UNESCO, Centre for Internet and Society (India), Adivasi Lives Matter, Digital Empowerment Foundation
- Other communities and conferences: Wikimedians from around the world, particularly during Wikimanias (2017–2019), Celtic Knot Conference (2018, 2019, 2025); Creative Commons Global Summits (2019 and 2020); Internet Governance Forum, Mozilla Festival 2021; National Geographic Citizen Science Workshop 2018;
- Grants, including a National Geographic Society grant and a Creative Commons grant, and support from Wikimedia Deutschland’s UNLOCK programme.
Competing interests
[edit | edit source]The author has no competing interest.
Ethics statement
[edit | edit source]This project draws direct/indirect learning from documentary films Gyani Maiya (2019), Mage Porob (2019) and Remosam (2019) that were made in collaboration respectively with the Kusunda community of Nepal, and Ho community and Bonda community of India. The participating individual members of these communities were interviewed with consent abided by the consent guidelines outlined in this project and the National Geographic Society release. Traditional community ethics were abided in all places while working together with indigenous groups and a high standard of moral and ethical standard was adhered to otherwise.
References
[edit | edit source]- ↑ UNESCO Ad Hoc Expert Group on Endangered Languages (in English). Paris: International Expert Meeting on UNESCO Programme Safeguarding of Endangered Languages, UNESCO. 2003. pp. 7–8. https://ich.unesco.org/doc/src/00120-EN.pdf.
- ↑ Avila, Eddie (2017). "How indigenous digital activists are leveraging the internet to revitalize their native languages". Linguapax Review 5: 80–89. http://www.linguapax.org/wp-content/uploads/2018/11/LinguapaxReview2017_web-1.pdf.
- ↑ Avila, Eddie (2021). "Technology in Language Revitalization: Rising Voices". In Olko, Justyna; Sallabank, Julia. Revitalizing Endangered Languages: A Practical Guide. Cambridge: Cambridge University Press. pp. 315–316. doi:10.1017/9781108641142.018. ISBN 9781108641142. https://www.cambridge.org/core/books/revitalizing-endangered-languages/technology-in-language-revitalization/9C4ED484CB915554C249941840999821.
- ↑ 4.0 4.1 4.2 4.3 4.4 Seyfeddinipur, Mandana; Rau, Felix (September 2020). "Keeping it real: Video data in language documentation and language archiving". Language Documentation & Conservation 14: 503–519. ISSN 1934-5275. http://hdl.handle.net/10125/24965.
- ↑ 5.0 5.1 Le Guen, Laila; Panigrahi, Subhashish (2021). "OpenSpeaks Accessibility". Wikimedia Deutschland. Archived from the original on 2021-11-22.
- ↑ 6.0 6.1 "OpenSpeaks Multimedia Toolkit". O Foundation. 2019-08-25. Retrieved 2026-03-18.
- ↑ 7.0 7.1 Wikiversity contributors (2021-11-22). "OpenSpeaks". Wikiversity. Retrieved 2026-03-18.
- ↑ 8.0 8.1 8.2 Panigrahi, Subhashish, Gomango, O., Pal, K. (2026), OpenSpeaks Archives: Citing Low-Resourced Language Oral History Multimedia, Wikimedia Foundation
- ↑ "International Conference Language Technologies for All (LT4All): Enabling Linguistic Diversity and Multilingualism Worldwide" (PDF). UNESCO. 2019-12-04. Retrieved 2026-03-18.
- ↑ Anderson, Gregory D. S.; Gomango, Opino (2016). "On the current status and state of Juray in the Sora-Juray cluster". In Ostler, Nicholas; Mohanty, Panchanan. FEL XX: Language Colonization and Endangerment: Long-term effects, echoes and reactions: Proceedings of the 20th FEL Conference 9–12 December 2016. Hungerford, England: FEL. pp. 103–109. ISBN 9780956021083.
- ↑ Panigrahi, Subhashish (2020-05-11). "Promoting coronavirus education through indigenous languages". Global Voices. Retrieved 2026-03-18.
- ↑ Silvertown, Jonathan (2009). "A new dawn for citizen science". Trends in Ecology & Evolution 24 (9): 467–471. doi:10.1016/j.tree.2009.03.017. ISSN 1872-8383.
- ↑ 13.0 13.1 Gebru, Timnit (2020). "Race and Gender". In Dubber, Markus D.; Pasquale, Frank; Das, Sunit. The Oxford Handbook of Ethics of AI. Oxford University Press. pp. 251–269. doi:10.1093/oxfordhb/9780190067397.013.16. ISBN 9780190067397.
- ↑ Panigrahi, S., Murmu, R. A. B., Baskey, F., Murmu, J. S. (2021), ᱥᱟᱛᱷᱟᱢ ᱑: ᱟᱝᱜᱚᱪ, ᱟᱝᱜᱚᱪ ᱟᱹᱭᱫᱟᱹᱨᱤᱠᱚ ᱟᱨ ᱟᱹᱭᱫᱟᱹᱨᱤ ᱞᱟᱭᱥᱮᱱᱥ, Wikiversity
- ↑ 15.0 15.1 15.2 15.3 15.4 Panigrahi, Subhashish; Tudu, Ramjit (2020-11-28). "We're updating OpenSpeaks and we'd love to hear from you!". O Foundation. Retrieved 2026-03-18.
- ↑ "OpenSpeaks/Archives". Meta-Wiki. Retrieved 2026-03-18.
- ↑ Panigrahi, Subhashish (2025-05-21). Open Speaks Archives: Learning from A Pilot to Enrich Wikipedia with Citable Oral History in Low-Resourced Languages. Wiki Workshop 2025. Online: Wikimedia Foundation.
- ↑ Lovo, Etivina; Woodward, Lynn; Larkins, Sarah; Preston, Robyn; Baba, Unaisi Nabobo (2021-10-09). "Indigenous knowledge around the ethics of human research from the Oceania region: A scoping literature review". Philosophy, Ethics, and Humanities in Medicine 16 (1). doi:10.1186/s13010-021-00108-8. ISSN 1747-5341. http://dx.doi.org/10.1186/s13010-021-00108-8.
- ↑ 19.0 19.1 "OpenSpeaks OERs". OpenSpeaks. Retrieved 2026-03-18.
- ↑ Brown, Penelope; Sicoli, Mark A.; Le Guen, Olivier (2021-10). "Cross-speaker repetition and epistemic stance in Tzeltal, Yucatec, and Zapotec conversations". Journal of Pragmatics 183: 256–272. doi:10.1016/j.pragma.2021.07.005. ISSN 0378-2166. http://dx.doi.org/10.1016/j.pragma.2021.07.005.
- ↑ "OpenSpeaks (Santali)". OpenSpeaks. Retrieved 2026-03-18.
- ↑ "OpenSpeaks events". Meta-Wiki. Retrieved 2026-03-18.
- ↑ Littell, Patrick; Kazantseva, Anna; Kuhn, Roland; Pine, Aidan; Arppe, Antti; Cox, Christopher; Junker, Marie-Odile (2018). "Indigenous language technologies in Canada: Assessment, challenges, and successes". Proceedings of the 27th International Conference on Computational Linguistics. Santa Fe, New Mexico: International Committee on Computational Linguistics. pp. 2620–2632. ISBN 978-1-948087-50-6. https://aclanthology.org/C18-1222.pdf.
- ↑ Kung, Susan; Smythe, Susan; Pojman, Elena; Niwagaba, Alicia (2020). "Archiving for the Future: Simple Steps for Archiving Language Documentation Collections". Retrieved 2026-03-18.