Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

3 Commits
 
 
 
 
 
 
 
 

Repository files navigation

Android Bench Community Results & Discussions

Welcome to the Android Bench Community Hub!

Android Bench is a benchmark for evaluating Large Language Models (LLMs) on tasks specific to Android development.

This repository serves as the central hub for the community to discuss evaluations, ask questions, share insights, and submit new model evaluations.

Evaluation Results on Harbor

The official community evaluation dataset and run results are hosted on Harbor Hub.

You can browse, download, and analyze trial logs, evaluation metrics, and model outputs directly on the Harbor Hub.

Community & Discussions

We encourage all users, researchers, and developers to engage with the project:

  • Ask Questions & Seek Help: Have a question about running benchmarks or interpreting results? Start a conversation on GitHub Discussions.
  • Propose & Discuss Evaluations: Share insights on model performance or discuss evaluation methodology.
  • Report Issues: Found an issue with evaluation data or repository documentation? Open an Issue.

Publishing & Sharing Evaluation Results

Check out our CONTRIBUTING.md guide for instructions on:

  1. Running evaluations locally using the Harbor CLI (harbor run -d android-bench/android-bench).
  2. Publishing evaluation runs to the Harbor Dataset.
  3. Sharing links to your published Harbor runs in GitHub Discussions to engage with the community.

Resources & Links

License

We are using Apache 2.0 License - see LICENSE for more information.

About

Android Bench Community model evaluation results

Resources

Contributing

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors