A benchmark for AGI
I have crazy ideas.
I kept trying...
• make good models
• use models in a good way
• try prompt engineering
• everything
But none of it made me feel like it was AGI.
A lot of work was poured into this...
It's too expensive.
Research.
Coding.
Training.
Fine-tuning.
Prompting.
All of it.
Not in mixed ways.
Not on free trials.
It's never going to work.
The reason for this dataset is one:
Let others make AGI.
It's still not perfect.
I don't say it's accurate.
But it is the farthest I could go...
see you soon.