M1A1 - Data Curation and Listing
Inventory of available open-source datasets for Wolof and Bambara across different categories: (1) English/French-Bambara translation pairs, (2) English/French-Wolof translation pairs, (3) Audio-transcription pairs, (4): Raw audio datasets without annotation. Establish data annot…
M1A2 - Synthetic Data Generation for Translation
The goal here will be to leverage our existing Wolof and Bambara LLM Translator to generate synthetic data at scale from open-source English corpora. The goal is to create small translator LLMs that are distilled from the larger models. Note : Distillation is when one a large mod…
M1A3 - Create Text & speech embedding eval dataset
A critical gap we address is the lack of evaluation benchmarks and high-quality datasets for African languages. Text and speech embedding models are used in RAG systems to find relevant documents based on user queries. To properly evaluate these models, we need benchmark datasets…
M2A1 - LLM Optimization for On-device/Offline Use
The idea here is to focus on small and frugal models, which are the most useful in african context due to internet connectivity limitations. Our goal would be to optimize models (like Oolel-Small) for size and efficiency for on-device deployment scenarios by distilling a small mo…
M2A2 - HuggingFace Space Deployment
Live interactive demo where users can test Wolof translation model and provide feedback, enabling community engagement and informal validation
M3A1 - Define the visual identity for the Soynade
Complete brand identity established for Soynade including logo, color palette, typography, and illustration style for consistent communication and marketing materials.
M3A2 - Design a website for Soynade
Fully designed and functional Soynade website outlining company mission, products, and contact information to establish online presence.
M4A1 - Market Research and Customer Discovery
Since we already have models performing well in translation, transcription, and analyzing English/French documents in Wolof or Bambara, we will develop a viable business plan around these capabilities, initially focusing on a streamlined SaaS offering. This activity will help val…
M4A2 - SaaS Concept Development
Defined SaaS product concept with specific features, target market, and value proposition.
M5A1 - Attend technical product dev session
The members of Soynade team will be joining a UNICEF Ventures workshop to meet the fellow portfolio companies and attend technical work sessions.
M5A2 - Revision of Q1 work plan
Revise the current workplan, by adding what have been achieved and what need to be adapted
M5A3 - Creation of company Karma profile
The creation of the Soynade profile and project file on Grantee Accountability Protocol (Karma)
M6A1 - Communications branding on Social media
We plan to post one image and caption in our linkedin and social media to announce a new release of dataset, model or for announcing a new project.
M6A2 - Communications and Branding (BlogPost)
We plan to write blogposts in order to share our vision of AI and open-source.
M7A1 - Establish licensing strategy
Clear open source licensing strategy established and OSI-approved licenses applied to all public repositories, ensuring legal compliance and community clarity
M7A2 - Create a Soynade project charter
Establish project charter defining Soynade's open source project vision, mission, community guidelines, and intellectual property strategy.
M7A3 - Ensure repository contains clear README
Create READMEs (in English) for all public repositories. READMEs should include: overview of specific repo, developer environment instructions (i.e. how to set software up), note about how repo connects into overall product, list of any Open Source software used to create product…
M7A4 - Public documentation
Create a public Open Source documentation github web page ensuring up-to-date technical documentation accessible to community for all project specific repositories
M7A5 - Establish an Open Source QA process
Established QA processes appropriate to project types (data validation, model evaluation, etc.) ensuring code and model quality before public release.
M7A6 - Establish Code of Conduct
Identify a Code of Conduct for any public Open Source repositories. Upload it to public source code repositories. Create internal documentation for how to respond to a Code of Conduct report, if one were to be made
M7A7 - Standardize Pull Request Workflow
Standardize pull request workflow adopted across all repositories ensuring code review, quality control, and collaborative development practices.
M8A1- Data privacy and security awareness session
Team members trained on data privacy and security best practices
M8A2 - Complete the Privacy Impact Assessment
Privacy Impact Assessment using UNICEF template documenting data processing practices and identifying any privacy risks for Soynade products.
Q2M1A1:Protocol for SpeechLLM data eval collection
One of our goals is to open-source high-quality evaluation datasets for the open-source and academic community working on West-African languages. Because one of the difficulties is the lack of a good evaluation dataset that is a standard for everyone. For this activity, our goal…
Q2M1A2: Speech Data Annotation for Evaluation
Curated and annotated copyright-free speech dataset specifically for evaluating Wolof Speech LLM capabilities
Q2M2A1: Model Evaluation
Comprehensive benchmark results for Wolof Speech LLM across multiple capabilities (transcription, translation, etc), documenting model performance and limitations to inform future development.
Q2M3A1: Model Evaluation
Rigorous assessment of Wolof text and speech embedding models on document retrieval tasks, quantifying accuracy and efficiency to validate readiness for production RAG applications.
Q2M8A1: Develop privacy agreements
Develop documents and/or privacy agreements to include privacy-specific language in the Terms of Use, vendor contracts, user consent notice, and any other agreement that may be applicable.
Q2M8A2: Develop a privacy policy
A comprehensive and transparent privacy policy is created, outlining the types of data collected (user inputs, personal information), the purpose of its use (service improvement, analytics), storage duration, and security measures.
Q2M8A3:Develop a data retention policy
A formal data retention policy is established that specifies how long different types of user and system data are stored, along with the procedures for secure deletion after that period.
Q2M6A1: OSI-approved license public source code
All public GitHub repositories for Soynade Research projects will have an OSI-approved open-source license (e.g., MIT, Apache 2.0, or GPL v3) formally applied, making the code legally open source and reusable by the community.
Q2M6A2: Contributing guidelines on OS repositories
CONTRIBUTING.md files established in all repositories with clear instructions on how external developers can contribute code, report issues, and submit pull requests.
Q2M6A3: Create tickets/issues for known bugs
Create public tickets/issues that correspond to planned features and known bugs/problems with Open Source repositories.
Q2M6A4:Project management board to track progress
Use a public project management board to track progress on public tickets/issues (e.g. Taiga, GitHub/GitLab Projects, JIRA, Trello, or similar).
Q2M6A5: User/Developer Documentation
Add either developer or user documentation to the Open Source documentation site. (Hint: Developer docs often include API docs, architecture or system state diagrams, or deployment guides.)
Q2M6A6: Open Source quality assurance
Advance Open Source quality assurance. Target 15% code coverage for unit tests, if applicable.
Q2M5A1: Revision of the work plan
Revise the current workplan, by adding what have been achieved and what need to be adapted
Q2M5A2: Milestone updates on KARMA
Update project milestones, activities, impact and/or attestations on Grantee Accountability Protocol (Karma)
Q2M5A3: Mid-term narrative report
Write a narrative report about Soynade’s vision for open-source and how this can benefit under-represented languages in AI, as well as how UNICEF funding helped us make this vision a reality.”
Q2M4A1: Saas Development
Continue developing the SaaS concept defined in Q1 based on the product and functionalities identified for target enterprises. In this milestone, we develop a function user interface for SaaS (developer console dashboard) enabling users to interact with Wolof language models thro…
Q2M4A2: Speech/Text LLM Integration into the SaaS
Integrate the Wolof speech LLM into the SaaS platform to enable voice-based interactions (e.g financial services use cases that accepts and understands voice commands).
Q2M4A3:Business Outreach and Funding Opportunities
Established visibility in relevant ecosystems through pitch materials and identified partnership/funding opportunities. Soynade positioned for potential collaborations and additional funding.
2025 - Data and Trust Cohort
Soynade Research develops open-source AI technologies including LLMs, translation tools, and speech systems for underserved African languages. Selected for the UNICEF Venture Fund program, we're building language technologies that serve millions of speakers across the continent.…