You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
We want to create evals for judging how well our LLMs perform on inline completions (ctrl+K), whole-file edits (ctrl+L), and autocomplete (tab).
A good starting task is to find an open-source eval for judging ctrl+L file completions (where the LLM rewrites a file given instructions). If you know of a high-quality data set for any of these tasks, we'd love to hear about it.
The text was updated successfully, but these errors were encountered:
We want to create evals for judging how well our LLMs perform on inline completions (ctrl+K), whole-file edits (ctrl+L), and autocomplete (tab).
A good starting task is to find an open-source eval for judging ctrl+L file completions (where the LLM rewrites a file given instructions). If you know of a high-quality data set for any of these tasks, we'd love to hear about it.
The text was updated successfully, but these errors were encountered: