-
Notifications
You must be signed in to change notification settings - Fork 2
Expand file tree
/
Copy pathabout.html
More file actions
129 lines (121 loc) · 5.35 KB
/
Copy pathabout.html
File metadata and controls
129 lines (121 loc) · 5.35 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
---
layout: default
title: About
permalink: /about/
description: >-
What EthioNLP is, how it started, and how it works.
---
{%- assign cat = site.data.generated.catalog -%}
{%- assign prog = site.data.generated.progress -%}
{%- assign hf = site.data.generated.huggingface -%}
{%- assign members = site.members -%}
<div class="page-head">
<div class="wrap wrap--wide">
<p class="eyebrow">About</p>
<h1 class="page-head__title">A research community for Ethiopian languages</h1>
<p class="lede">
EthioNLP began at COLING 2018 in Santa Fe, among researchers from Addis
Ababa University, the University of Hamburg and the University of Trento
who proposed a shared community for Ethiopian language technology rather
than separate individual efforts.
</p>
<ul class="statline">
<li><strong>80+</strong> languages in Ethiopia</li>
<li><strong>{{ cat.totals.with_artifacts }}</strong> with an open model or dataset from us</li>
<li><strong>{{ members | size }}</strong> people listed</li>
<li><strong>2018</strong> founded</li>
</ul>
</div>
</div>
<section class="band">
<div class="wrap wrap--wide">
<div class="prose">
<h2>Why it exists</h2>
<p>
NLP work on Ethiopian languages has been carried out in many places
without formal communication between the researchers involved.
Groups often worked on the same problems, rebuilt similar corpora,
and published in a literature that had not been collected in one place.
</p>
<p>
The community is organised around collaboration and training as well as
output. Members co-author, review each other's work, supervise theses and
run workshops. Datasets and models follow from that, and training a
researcher to run their own experiments has value beyond a single
release.
</p>
<p>
The same applies to anyone who contributes language data. A speaker,
annotator or collector is a collaborator on the work, named in it where
they wish to be, and paid fairly where a project has funding. The
expectations are set out in the
<a href="{{ '/standards/' | relative_url }}">community standards</a>.
</p>
<p>
Ethiopia is multinational and multilingual, with more than 80
languages. The catalogue lists
<a href="{{ '/languages/#sources' | relative_url }}">{{ cat.totals.languages }} entries</a>,
following Glottolog, which also includes extinct, pidgin and signed
varieties.
These languages are under-represented in language technology. Of the four
Ethiopian languages with enough indexed NLP papers to measure,
<strong>{{ prog.totals.top_language_share }}% of the work is on
{{ prog.totals.top_language }} alone</strong>. Many of the languages in
the catalogue have no corpus, no baseline and no paper to cite.
</p>
<h2>How it works</h2>
<p>
Everything here is public and regenerated from public sources. Models and
datasets come from the Hugging Face API, code from GitHub, publications
from OpenAlex, Semantic Scholar, DBLP and the ACL Anthology, and the
language catalogue from Glottolog and Wikidata. Profiles are not
maintained by hand; a member supplies a few public identifiers once, and
the site updates from them.
</p>
<p>
Two distinctions are enforced. Work <a href="{{ '/ecosystem/' | relative_url }}">built
by this community</a> is kept apart from
<a href="{{ '/ecosystem/related/' | relative_url }}">Ethiopian-language work by
others</a>, and work built <em>for</em> an Ethiopian language is kept
apart from massively multilingual work that lists one among many.
</p>
<p>
Members are listed only after they confirm it. A Hugging Face
organisation has more people in it than this site lists; only the
{{ members | size }} who have agreed to appear are shown.
</p>
<h2>Part of AfricaNLP</h2>
<p>
EthioNLP takes part in the wider African NLP community, including
Masakhane, the AfricaNLP workshop series and the Deep Learning
Indaba. The
<a href="{{ '/progress/' | relative_url }}">progress dashboard</a> tracks
Ethiopian-language research against African-language research as a
whole.
</p>
<h2>Taking part</h2>
<p>
There is no vetting step and no fee. Students are welcome; many of the
open topics are the size of a first thesis rather than a longer research
programme. See <a href="{{ '/join/' | relative_url }}">how to join</a>,
the <a href="{{ '/mentorship/' | relative_url }}">open thesis topics</a>,
or <a href="{{ '/teams/' | relative_url }}">how a lab registers</a>.
</p>
</div>
</div>
</section>
<section class="band band--alt">
<div class="wrap wrap--wide">
<div class="section-head">
<h2>What we work on</h2>
<p class="section-head__note">Six tracks, described in full on the research page.</p>
</div>
<div class="grid grid--3">
{%- for area in site.data.focus_areas %}
<a class="card card--link" href="{{ '/research/' | relative_url }}#{{ area.key }}">
<h3 class="card__title">{{ area.title }}</h3>
</a>
{%- endfor %}
</div>
</div>
</section>