Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Controlling the generation for code quality will be extremely hard.

The only thing I see is that they could filter their dataset so that some bad proxy of code quality is taken into account, something like the number of stars (which is clearly a terrible metric, tell me if you think of something else).

The idea would maybe start by training on all of the subset of Github it is ethical and legal to train on, and then filter down to higher code quality towards the end.

Controlling for the time at which the code is emitted would be easier. Something like, retrieving similar contexts, and guiding the model to be more similar to the recent code if there is similar recent code that exists. I'm not sure exactly of how this would be done, but I can see it working.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: