The Story Behind BENI

Ann Naser Nabil · July 2026

It started with a simple question: Why is all economic NLP research about English?

84% of sentiment-based forecasting focuses on developed economies. Bangla, the 7th most spoken language in the world? Almost nothing.

The Dataset

BENI v1.0 is a Bangla Economic Narrative Index dataset. 620K+ articles across 10 languages from the Global South. Published on HuggingFace and arXiv.

What I Learned

  1. Low-resource NLP is hard because the data doesn't exist.
  2. Economic narratives matter — they shape policy, markets, perceptions.
  3. Building the dataset IS the research.

Read the paper: arXiv:2606.10225