Trained on a different random sampling of the same datasets used by loyal-piano-m7, then with cDPO on a blend of RLHF datasets.
Several intermediate checkpoints (of cDPO training) are on branches.
Uses the Alpaca prompt format.
- Downloads last month
- 737
This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social
visibility and check back later, or deploy to Inference Endpoints (dedicated)
instead.