Welcome to the Android Bench Community Hub!
Android Bench is a benchmark for evaluating Large Language Models (LLMs) on tasks specific to Android development.
This repository serves as the central hub for the community to discuss evaluations, ask questions, share insights, and submit new model evaluations.
The official community evaluation dataset and run results are hosted on Harbor Hub.
You can browse, download, and analyze trial logs, evaluation metrics, and model outputs directly on the Harbor Hub.
We encourage all users, researchers, and developers to engage with the project:
- Ask Questions & Seek Help: Have a question about running benchmarks or interpreting results? Start a conversation on GitHub Discussions.
- Propose & Discuss Evaluations: Share insights on model performance or discuss evaluation methodology.
- Report Issues: Found an issue with evaluation data or repository documentation? Open an Issue.
Check out our CONTRIBUTING.md guide for instructions on:
- Running evaluations locally using the Harbor CLI (
harbor run -d android-bench/android-bench). - Publishing evaluation runs to the Harbor Dataset.
- Sharing links to your published Harbor runs in GitHub Discussions to engage with the community.
- Android Bench Website: developer.android.com/bench
- Harbor Dataset: hub.harborframework.com/datasets/android-bench/android-bench/latest
- Harbor Documentation: harborframework.com/docs
- Community Tasks Repository: android-bench/community-dataset
We are using Apache 2.0 License - see LICENSE for more information.