पाठ 19 / 27
Tokens, Pricing और लागत का अनुमान
Ship करने से पहले usage संख्याओं को पैसे में बदलें।
Input और output अलग क़ीमत के हैं
Providers प्रति दस लाख input tokens और प्रति दस लाख output tokens क़ीमत लगाते हैं, output आम तौर पर कई गुना महँगा, और बड़े या नए मॉडल छोटे से महँगे। कुछ सुविधाओं (cached prompt prefixes, batch processing, images, tool परिभाषाएँ, लंबा context) की अपनी दरें हैं। इसलिए call की लागत (input_tokens x input_price + output_tokens x output_price) / 1,000,000 है। Launch से पहले अनुमान लगाएँ: प्रति अनुरोध औसत input और output tokens गुणा अपेक्षित मासिक requests। याद रखें कि tool परिभाषाएँ, system prompts और बातचीत का इतिहास हर अनुरोध पर input tokens गिने जाते हैं। क़ीमतें बदलती हैं, इसलिए provider का मौजूदा pricing पृष्ठ पढ़ें; नीचे के आँकड़े गढ़े उदाहरण हैं, असली क़ीमतें नहीं।
जानें हर call की क्या लागत है
लागत tokens गुणा क़ीमत है; prompts छोटे करें, prefixes cache करें, सही मॉडल चुनें और usage ट्रैक करें।
लागत कैलकुलेटर, चलाकर
मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। गढ़ी क़ीमतों के साथ वही 1,200-अंदर / 150-बाहर call छोटे मॉडल पर लगभग $0.0005 और बड़े पर $0.0059 की है; बड़े मॉडल पर लंबा 8,000-token prompt $0.033 का। अंतिम पंक्ति दिखाती है कि प्रति call की छोटी लागत 10,000 calls में $4.88 कैसे बनती है।
PRICES = {"small-model": (0.25, 1.25), "large-model": (3.00, 15.00)} # example USD per million tokens, NOT real prices
def cost(model, input_tokens, output_tokens):
pin, pout = PRICES[model]
return (input_tokens * pin + output_tokens * pout) / 1_000_000
calls = [("small-model", 1200, 150), ("large-model", 1200, 150), ("large-model", 8000, 600)]
total = 0.0
for model, i, o in calls:
c = cost(model, i, o); total += c
print(f"{model:12} in={i:5} out={o:4} -> ${c:.6f}")
print("total: $%.6f" % total)
print("10,000 calls like the first one: $%.2f" % (10_000 * cost("small-model", 1200, 150)))
Output:
small-model in= 1200 out= 150 -> $0.000487 large-model in= 1200 out= 150 -> $0.005850 large-model in= 8000 out= 600 -> $0.033000 total: $0.039338 10,000 calls like the first one: $4.88
भेजने से पहले tokens गिनें
Prompts का आकार जाँचने और बहुत बड़े इनपुट जल्दी अस्वीकार करने को provider की token गिनती उपयोग करें।
त्वरित जाँच: हर अनुरोध पर input tokens में क्या गिना जाता है?
- सिर्फ़ उत्तर
- सिर्फ़ user का नवीनतम शब्द
- कुछ नहीं
- System prompt, tool परिभाषाएँ और बातचीत का इतिहास
Answer
System prompt, tool परिभाषाएँ और बातचीत का इतिहास — आप जो भी भेजते हैं बिल होता है, दोहराया संदर्भ भी।