🧠
Model

Appworld Leaderboard

by StonyBrookNLP stonybrooknlp/appworld-leaderboard
Free2AITools Nexus Index
42.5
S: Semantic 50

Query-time baseline · scored live at search

A: Authority 49
P: Popularity 46
R: Recency 64
Q: Quality 70
Tech Context
Vital Performance —

Task categories from upstream metadata

πŸ’¬Chat & Dialogue

Technical Constraints

Experimental / High Latency
Low FNI signal 42.5 FNI Score
Tiny - Params
- Context
0 Downloads
Model Information Summary
Entity Passport
Registry ID stonybrooknlp/appworld-leaderboard
Provider github
πŸ“œ

Cite this model

Academic & Research Attribution

BibTeX
@misc{stonybrooknlp_appworld_leaderboard,
  author = {StonyBrookNLP},
  title = {Appworld Leaderboard Model},
  year = {2024},
  howpublished = {\url{https://github.com/StonyBrookNLP/appworld-leaderboard}},
  note = {Accessed via Free2AITools.}
}
APA Style
StonyBrookNLP. (2024). Appworld Leaderboard [Model]. Free2AITools. https://github.com/StonyBrookNLP/appworld-leaderboard

πŸ”¬Technical Deep Dive

Full Specifications [+]

Quick Commands

πŸ™ Git Clone
git clone https://github.com/StonyBrookNLP/appworld-leaderboard

βš–οΈ Free2AITools Nexus Index V2.0

Semantic (S) 50

Query-time baseline · scored live at search

Authority (A) 49
Popularity (P) 46
Recency (R) 64
Quality (Q) 70

πŸ’¬ Index Insight

FNI V2.0 for Appworld Leaderboard: Authority (A:49), Popularity (P:46), Recency (R:64), Quality (Q:70). Semantic (S) is a query-time baseline scored live at search.

Free2AITools Nexus Index

Data Sources / Provenance

Open data Updated: Live data
---

πŸš€ What's Next?

Technical Deep Dive

AppWorld Leaderboard

This is the leaderboard repository of the benchmark proposed in AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents (ACL 2024).

The project's main repository is here, and leaderboard UI is here. This repository stores bundled (encrypted) experiment outputs from participanting models (including our baselines) and the raw leaderboard data JSON which is dynamically rendered in the UI. You can use this repository to:

  1. Download and locally view experiment outputs from other participanting methods.
  2. Submit your own agent's experiment outputs to be included on the leaderboard via a PR.

For both cases, you first need to

  1. Install appworld: pip install appworld && appworld install.
  2. Install Git LFS and clone the appworld-leaderboard repository.

Submit Your Agent's Outputs

First, pack your agent's test_normal and test_challenge experiment outputs:

::Click:: Experiment outputs refresher

Your experiment outputs are located in ./experiments/outputs/{experiment_name} relative to the APPWORLD_ROOT, which as we discussed earlier, defaults to ., but can be configured by passing APPWORLD_ROOT environment variable or --root in CLI.


For the leaderboard, experiment names must be alphanumeric with optional hyphens and underscores, and they must end with the dataset name, i.e., _test_normal or _test_challenge, e.g., react_gpt4o_test_normal. You should have two experiment outputs, one for each dataset. The prefix portion of their names must be the same, e.g., react_gpt4o_test_normal and react_gpt4o_test_challenge. Rename the directory accordingly if necessary.

Now, pack the two experiments, individually, with the following commands, and same metadata.

bash
appworld pack {test_normal_experiment_n

⚠️ Incomplete Data

Some information about this model is not available. Use with Caution - Verify details from the original source before relying on this data.

View Original Source β†’

πŸ“ Limitations & Considerations

  • β€’ Benchmark scores may vary based on evaluation methodology and hardware configuration.
  • β€’ VRAM requirements are estimates; actual usage depends on quantization and batch size.
  • β€’ FNI scores are relative rankings and may change as new models are added.
  • ⚠ License Unknown: Verify licensing terms before commercial use.

Social Proof

GitHub Repository
6Stars
11Forks
πŸ”„ Updated daily

Source summary: Based on GitHub metadata. Not a recommendation.

πŸ“Š FNI Methodology πŸ“š Knowledge Baseℹ️ Verify with original source

πŸ›‘οΈ Model Transparency Report

Technical metadata sourced from upstream repositories.

Open Metadata

πŸ†” Identity & Source

id
gh-model--stonybrooknlp--appworld-leaderboard
slug
stonybrooknlp--appworld-leaderboard
source
github
author
StonyBrookNLP
tags
acl-2024, ai-agents, ai-apis, ai-assistants, ai-environment, ai-planning, autonomous-agents, coding-agents, function-calling, interactive-coding, llm, llm-agents, nlp-datasets, nlp-machine, tool-usage, ai, python

βš™οΈ Technical Specs

pipeline tag
text-generation

πŸ“Š Engagement & Metrics

downloads
0
stars
6
forks
11

Data indexed from public sources. Updated daily.