GraphQL is a query language for your API plus a type system describing what can be asked. One endpoint accepts a document, executes it against a schema, and returns exactly the requested fields — moving work from the client to the server, and from runtime into schema design.
Three problems push teams toward it. Overfetching: a phone list needs a title and a thumbnail, and GET /books sends forty fields per row because a desktop screen once needed them. Underfetching: a book page takes /books/42, then /authors/7, then /books/42/reviews, three sequential round trips because each depends on the last. Version drift: every client wants a different shape, so the API grows ?fields=, ?include=, /v2 and a -mobile variant. Sparse fieldsets (Filtering and Sorting) solve part of this; GraphQL solves all of it, from a schema that doubles as an introspectable contract.
What you pay is concrete. Every request is a POST with a different body, so HTTP caching — Cache-Control, ETags, ETags and Conditional Requests — gives way to server-side caching and persisted queries. Rate limiting by request count becomes meaningless, since one request can ask for ten thousand nodes. A failed field comes back as 200 with an errors array, breaking dashboard rules that count 5xx responses. And a naive resolver turns one query into hundreds of database calls (Resolvers and N+1).
So pick GraphQL when several independent clients — web, iOS, Android, partners — consume one domain graph and diverge faster than you can ship endpoints. Stay with REST for a single first-party client, for responses a CDN should cache, and wherever uploads dominate. A mixed API is normal and usually right.