Context window
The maximum tokens a model can consider at once — prompt plus output. Exceed it and the request fails or history silently truncates, which is how chatbots 'forget' early instructions.
The maximum tokens a model can consider at once — prompt plus output. Exceed it and the request fails or history silently truncates, which is how chatbots 'forget' early instructions.