Second, it has the problem of non-stoping response.
I see non-stop response as a generalization problem because normally every training sample is not of infinite length.
Targeted supervised fine-tuning should work, as long as you have enough samples. However, supervised fine-tuning is not good for generalization.
Second, it has the problem of non-stoping response.