पाठ 19 / 27

Tokens, Pricing और लागत का अनुमान

Ship करने से पहले usage संख्याओं को पैसे में बदलें।

Input और output अलग क़ीमत के हैं

Providers प्रति दस लाख input tokens और प्रति दस लाख output tokens क़ीमत लगाते हैं, output आम तौर पर कई गुना महँगा, और बड़े या नए मॉडल छोटे से महँगे। कुछ सुविधाओं (cached prompt prefixes, batch processing, images, tool परिभाषाएँ, लंबा context) की अपनी दरें हैं। इसलिए call की लागत (input_tokens x input_price + output_tokens x output_price) / 1,000,000 है। Launch से पहले अनुमान लगाएँ: प्रति अनुरोध औसत input और output tokens गुणा अपेक्षित मासिक requests। याद रखें कि tool परिभाषाएँ, system prompts और बातचीत का इतिहास हर अनुरोध पर input tokens गिने जाते हैं। क़ीमतें बदलती हैं, इसलिए provider का मौजूदा pricing पृष्ठ पढ़ें; नीचे के आँकड़े गढ़े उदाहरण हैं, असली क़ीमतें नहीं।

जानें हर call की क्या लागत है

लागत tokens गुणा क़ीमत है; prompts छोटे करें, prefixes cache करें, सही मॉडल चुनें और usage ट्रैक करें।

तीन उत्तोलक: गिनें, cache करें, चुनें।
चित्र 6.1 — गिनें, cache करें और चुनें।

लागत कैलकुलेटर, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। गढ़ी क़ीमतों के साथ वही 1,200-अंदर / 150-बाहर call छोटे मॉडल पर लगभग $0.0005 और बड़े पर $0.0059 की है; बड़े मॉडल पर लंबा 8,000-token prompt $0.033 का। अंतिम पंक्ति दिखाती है कि प्रति call की छोटी लागत 10,000 calls में $4.88 कैसे बनती है।

PRICES = {"small-model": (0.25, 1.25), "large-model": (3.00, 15.00)}      # example USD per million tokens, NOT real prices

def cost(model, input_tokens, output_tokens):
    pin, pout = PRICES[model]
    return (input_tokens * pin + output_tokens * pout) / 1_000_000

calls = [("small-model", 1200, 150), ("large-model", 1200, 150), ("large-model", 8000, 600)]
total = 0.0
for model, i, o in calls:
    c = cost(model, i, o); total += c
    print(f"{model:12} in={i:5} out={o:4} -> ${c:.6f}")
print("total: $%.6f" % total)
print("10,000 calls like the first one: $%.2f" % (10_000 * cost("small-model", 1200, 150)))

Output:

small-model  in= 1200 out= 150 -> $0.000487
large-model  in= 1200 out= 150 -> $0.005850
large-model  in= 8000 out= 600 -> $0.033000
total: $0.039338
10,000 calls like the first one: $4.88

भेजने से पहले tokens गिनें

Prompts का आकार जाँचने और बहुत बड़े इनपुट जल्दी अस्वीकार करने को provider की token गिनती उपयोग करें।

त्वरित जाँच: हर अनुरोध पर input tokens में क्या गिना जाता है?

  • सिर्फ़ उत्तर
  • सिर्फ़ user का नवीनतम शब्द
  • कुछ नहीं
  • System prompt, tool परिभाषाएँ और बातचीत का इतिहास
Answer

System prompt, tool परिभाषाएँ और बातचीत का इतिहास — आप जो भी भेजते हैं बिल होता है, दोहराया संदर्भ भी।