Hi, I am testing the Inference API with different models to rewrite texts. But no matter which model I choose, only about 2 - 3 sentences are returned as response. How can this be adjusted?
A few things to check:
-
Make sure you’re setting
max_new_tokens(or the equivalent generation parameter) high enough. The default value can be quite small. -
Verify that the model you’re using supports long-form text generation.
-
Check whether your input plus output exceeds the model’s context window.
-
If you’re using the Inference API, inspect the full response to see if it indicates the generation stopped because it reached a token limit (
finish_reasonor similar).
If you can share the model name and your request payload, it will be easier to identify the cause.