Skip to content
Salesforce 19 min read 14 sections

Batch Apex in Salesforce, explained with examples

Every org reaches a day when one transaction cannot touch every record that needs touching. A data migration leaves two lakh Accounts with a wrong Rating. A nightly job has to recalculate a field on every Opportunity closed in the last year. Synchronous Apex gives you 50,000 query rows and 10,000 DM

RL

RizeX Labs

RizeX Labs

Published

Last updated

Related programme

Salesforce Training

Admin, Apex, Lightning Web Components and integrations in a live batch, with an industry project, mock interviews and a network of 270 hiring companies.

See the curriculum

3000+ learners · 270+ hiring partners

Every org reaches a day when one transaction cannot touch every record that needs touching. A data migration leaves two lakh Accounts with a wrong Rating. A nightly job has to recalculate a field on every Opportunity closed in the last year. Synchronous Apex gives you 50,000 query rows and 10,000 DML rows in a single transaction, and that is the ceiling. Batch Apex is the framework that splits the work into many small transactions so the ceiling applies to each chunk instead of the whole job. This tutorial walks through the interface, the three methods, the scope parameter, state, callouts, chaining, scheduling, testing, and the failures that show up only in production.

Why Batch Apex exists and which limits it resets

Batch Apex takes one large set of records, breaks it into chunks, and runs your logic once per chunk. Each chunk is its own transaction with its own governor limits.

That last sentence is the whole point of the feature. Here is what resets on every chunk:

Limit Value per execute transaction
SOQL queries 100
Query rows returned 50,000
DML statements 150
DML rows 10,000
CPU time 60,000 ms (asynchronous)
Heap size 12 MB (asynchronous)
Callouts 100, if the class allows callouts

So a job with 5,000 chunks gets 5,000 separate allocations of all of the above. That is how one million records get processed without a single LimitException.

Two things do not reset. The number of records the start method can hand over is capped at 50 million when you use a QueryLocator. And the 24-hour asynchronous limits are org-wide, not per chunk: a maximum of 250,000 execute method invocations per rolling 24 hours (or the number of user licences multiplied by 200, whichever is higher), and a maximum of five batch jobs queued or active at the same time.

This is the first place candidates slip. Question 32 in most question banks asks for the difference between Queueable and Batch Apex, and the memorised answer is "Queueable is for chaining, Batch is for large data". The real difference is shape. Queueable runs your logic once, in one transaction, with one set of limits, and is the right tool when the work is a single unit that should not block the user. Batch runs your logic many times, once per chunk, with a fresh set of limits each time, and is the right tool when the work scales with record count. If you can state it that way, the follow-up questions get easier.

The Database.Batchable interface and its three methods

A batch class implements Database.Batchable<sObject> and must define exactly three methods.

public class MyBatch implements Database.Batchable<sObject> {

    // 1. Collect the records. Runs once.
    public Database.QueryLocator start(Database.BatchableContext bc) {
        return Database.getQueryLocator('SELECT Id FROM Account');
    }

    // 2. Process one chunk. Runs once per chunk.
    public void execute(Database.BatchableContext bc, List<Account> scope) {
        // your logic
    }

    // 3. Clean up, notify, chain. Runs once, after all chunks.
    public void finish(Database.BatchableContext bc) {
        // post-processing
    }
}

start runs once at the beginning and returns the full set of records to process, either as a Database.QueryLocator or as an Iterable<sObject>.

execute runs once for every chunk. The scope parameter is the list of records in that chunk. You can type it as List<sObject> or, more usefully, as the concrete type such as List<Account>, which saves you a cast.

finish runs once after the last chunk completes, in its own transaction. Use it for the completion email, for a summary record, or to start the next batch in a chain.

Database.BatchableContext carries the job identity. bc.getJobId() returns the AsyncApexJob Id, which is how you query the job's own status from inside finish.

One trainer point that is worth absorbing before the interview: candidates know the syntax but not the execution. They can write the three methods and then cannot say when each one runs, in which transaction, or what data it can see. Being clear that start runs once, execute runs N times, and finish runs once in a separate transaction is worth more than reciting the signatures.

A complete batch class that updates Accounts

Here is a job that reads every Account with revenue on it and sets the Rating field based on that revenue.

public class AccountRatingBatch implements Database.Batchable<sObject> {

    public Database.QueryLocator start(Database.BatchableContext bc) {
        return Database.getQueryLocator(
            'SELECT Id, AnnualRevenue, Rating ' +
            'FROM Account ' +
            'WHERE AnnualRevenue != NULL'
        );
    }

    public void execute(Database.BatchableContext bc, List<Account> scope) {

        List<Account> toUpdate = new List<Account>();

        for (Account acc : scope) {
            String newRating = acc.AnnualRevenue >= 10000000 ? 'Hot' : 'Warm';
            if (acc.Rating != newRating) {
                acc.Rating = newRating;
                toUpdate.add(acc);
            }
        }

        if (!toUpdate.isEmpty()) {
            // allOrNone = false, so one bad record does not roll back the chunk
            Database.SaveResult[] results = Database.update(toUpdate, false);

            for (Integer i = 0; i < results.size(); i++) {
                if (!results[i].isSuccess()) {
                    System.debug(LoggingLevel.ERROR,
                        'Failed ' + toUpdate[i].Id + ' : ' +
                        results[i].getErrors()[0].getMessage());
                }
            }
        }
    }

    public void finish(Database.BatchableContext bc) {
        AsyncApexJob job = [
            SELECT Id, Status, ExtendedStatus, NumberOfErrors,
                   JobItemsProcessed, TotalJobItems, CreatedBy.Email
            FROM AsyncApexJob
            WHERE Id = :bc.getJobId()
        ];
        System.debug('Job ' + job.Status + ' with ' + job.NumberOfErrors + ' errors');
    }
}

Three details in that code are what separate a working batch from a demo batch.

The DML happens once per chunk, outside the loop, on a collected list. Putting update acc; inside the for loop would burn one of your 150 DML statements per record and fail at record 151.

Database.update(list, false) allows partial success. With a plain update list;, one validation rule failure rolls back the entire chunk of 200 records. With allOrNone set to false, the good records save and you get a SaveResult per record telling you which ones failed.

The if (acc.Rating != newRating) check means records already holding the correct value are never written. On a large object that single line can cut the DML volume by ninety percent and stop triggers and flows from firing pointlessly.

QueryLocator or Iterable in the start method

The start method can return one of two things, and the choice sets the hard ceiling on your job.

Database.QueryLocator Iterable<sObject>
Maximum records 50 million 50,000
Source A SOQL query Any collection or custom Iterable
Query rows limit Bypassed Counts against the normal 50,000 rows
Typical use Almost every real job Records built in code, or non-SOQL sources

Database.getQueryLocator is special. The rows it returns do not count against the 50,000 query-rows governor limit, and that exemption is the entire reason Batch Apex can process one million records. It is also why the answer to "how do you process one million records" is not "use Batch Apex" but "use Batch Apex with a QueryLocator in start, because an Iterable caps you at 50,000 and you would fail at record 50,001".

// QueryLocator: up to 50 million records
public Database.QueryLocator start(Database.BatchableContext bc) {
    return Database.getQueryLocator([
        SELECT Id, AnnualRevenue, Rating
        FROM Account
        WHERE AnnualRevenue != NULL
    ]);
}
// Iterable: only when the records are not a straight SOQL result.
// Still limited to 50,000 rows.
public Iterable<sObject> start(Database.BatchableContext bc) {
    List<Account> records = buildListFromSomeOtherSource();
    return records;
}

Use Iterable when the record set cannot be expressed as one SOQL query, for example when you are paging through a web service response or filtering on something SOQL cannot express. Use QueryLocator everywhere else.

One restriction to remember: a QueryLocator query cannot use aggregate functions such as COUNT() or SUM(), and it cannot use GROUP BY. If you need aggregated data, query it inside execute or finish, not in start.

The scope parameter, the default of 200 and the maximum of 2,000

The second argument to Database.executeBatch is the scope, meaning how many records go into each chunk.

Database.executeBatch(new AccountRatingBatch());        // scope = 200 (default)
Database.executeBatch(new AccountRatingBatch(), 50);    // scope = 50
Database.executeBatch(new AccountRatingBatch(), 2000);  // scope = 2000 (maximum)

The default is 200. The maximum you can pass is 2,000. Anything higher is rejected.

Larger scope means fewer chunks and fewer transactions, which finishes faster and uses fewer of your 24-hour execute invocations. Smaller scope means fewer records per transaction, which means less CPU and less heap consumed per chunk.

Reducing the scope is the correct fix when a job fails with a CPU time limit or an Apex heap size error. If your execute does heavy string work, builds large maps, or updates a record whose trigger cascades into three other objects, 200 records may be too many for one 60-second transaction. Drop it to 50, or to 20, and the same job completes. The job runs longer overall, and that is an acceptable trade.

Reducing the scope is not the fix for a "non-selective query" error in start, because start runs as one query regardless of scope. That is an indexing problem, covered further down.

Database.Stateful and what actually survives

By default, every execute transaction gets a fresh instance of your class. Any counter you increment in chunk one is back to zero in chunk two. Implement Database.Stateful and the instance variables carry across.

public class AccountRatingBatch implements Database.Batchable<sObject>, Database.Stateful {

    // instance variable: SURVIVES between execute calls
    public Integer recordsProcessed = 0;
    public List<String> failedIds = new List<String>();

    // static variable: DOES NOT survive. Reset on every chunk.
    public static Integer staticCounter = 0;

    public Database.QueryLocator start(Database.BatchableContext bc) {
        return Database.getQueryLocator(
            'SELECT Id, AnnualRevenue, Rating FROM Account WHERE AnnualRevenue != NULL'
        );
    }

    public void execute(Database.BatchableContext bc, List<Account> scope) {
        for (Account acc : scope) {
            acc.Rating = acc.AnnualRevenue >= 10000000 ? 'Hot' : 'Warm';
        }

        Database.SaveResult[] results = Database.update(scope, false);

        for (Integer i = 0; i < results.size(); i++) {
            if (results[i].isSuccess()) {
                recordsProcessed++;          // kept across chunks
            } else {
                failedIds.add(scope[i].Id);  // kept across chunks
            }
        }

        staticCounter++;                     // always 1 at the end of each chunk
    }

    public void finish(Database.BatchableContext bc) {
        System.debug('Processed ' + recordsProcessed + ' records');
        System.debug('Failures: ' + failedIds.size());
    }
}

What survives: instance member variables of the batch class.

What does not survive: static variables, and anything cached in a static map by a helper class. Each chunk is a new transaction and statics are reset per transaction.

Two cautions. Everything you keep in an instance variable is serialised between transactions, so a list that grows to 50,000 entries will eventually break the heap limit in a later chunk. Keep aggregates and counters, not full record collections. And a class that keeps state cannot be assumed to process chunks in order, because Salesforce does not guarantee chunk ordering.

Database.Stateful is also what makes finish useful. Without it, finish has no idea what happened in the chunks and can only query AsyncApexJob for a count.

Database.AllowsCallouts for HTTP calls

A batch class cannot make an HTTP callout unless it declares that it will.

public class AccountSyncBatch implements Database.Batchable<sObject>,
                                         Database.AllowsCallouts,
                                         Database.Stateful {

    public Integer synced = 0;

    public Database.QueryLocator start(Database.BatchableContext bc) {
        return Database.getQueryLocator(
            'SELECT Id, Name, External_Id__c FROM Account WHERE Sync_Pending__c = true'
        );
    }

    public void execute(Database.BatchableContext bc, List<Account> scope) {

        List<Account> toUpdate = new List<Account>();

        for (Account acc : scope) {
            HttpRequest req = new HttpRequest();
            req.setEndpoint('callout:ERP_Named_Credential/accounts/' + acc.External_Id__c);
            req.setMethod('GET');
            req.setTimeout(20000);

            try {
                HttpResponse res = new Http().send(req);
                if (res.getStatusCode() == 200) {
                    acc.Sync_Pending__c = false;
                    toUpdate.add(acc);
                    synced++;
                }
            } catch (CalloutException e) {
                System.debug(LoggingLevel.ERROR, acc.Id + ' : ' + e.getMessage());
            }
        }

        if (!toUpdate.isEmpty()) {
            Database.update(toUpdate, false);
        }
    }

    public void finish(Database.BatchableContext bc) {
        System.debug('Synced ' + synced + ' accounts');
    }
}

Without Database.AllowsCallouts, the same code throws System.CalloutException: Callout from scheduled Apex not supported or a similar callout-not-permitted error at runtime, not at compile time.

The limit is 100 callouts per execute transaction. With the default scope of 200 and one callout per record, you exceed it at record 101. Either drop the scope to 100 or fewer, or batch the records into one composite request.

The other trap is ordering. All DML must come after your callouts in a given transaction. If you update records first and then attempt a callout in the same execute, you get "You have uncommitted work pending". Collect records into a list, finish the callouts, then do the DML at the end, exactly as the code above does.

Chaining one batch from finish

finish runs in its own transaction, which makes it the correct place to start the next job.

public void finish(Database.BatchableContext bc) {

    // do not chain during tests, or the test will not complete
    if (!Test.isRunningTest()) {
        Database.executeBatch(new ContactRatingBatch(), 200);
    }
}

Chaining is how you run dependent jobs in sequence. Update Accounts first, then roll the result down to Contacts, then write a summary. Running all three at once would race each other and hit the five-concurrent-jobs limit.

Guard the chain with Test.isRunningTest(). Chained batch jobs do not execute in a test context, and leaving the call unguarded produces confusing test failures.

Do not build an unconditional loop where a batch chains back into itself. Give it an exit condition, such as a stateful counter or a query that returns zero remaining records, or the job runs until it exhausts the 24-hour asynchronous limit.

Running a batch and scheduling one

From Developer Console, Execute Anonymous:

Id jobId = Database.executeBatch(new AccountRatingBatch(), 200);
System.debug('Job Id: ' + jobId);

Database.executeBatch returns the AsyncApexJob Id straight away. The job itself goes into the Apex flex queue and starts when a slot is free. To stop a running job, use System.abortJob(jobId).

For a scheduled run, write a small Schedulable class that does nothing except start the batch.

public class AccountRatingScheduler implements Schedulable {

    public void execute(SchedulableContext sc) {
        Database.executeBatch(new AccountRatingBatch(), 200);
    }
}

Schedule it with a cron expression. The Apex cron format has seven fields: seconds, minutes, hours, day of month, month, day of week, and an optional year.

// every day at 2:00 AM
String cron = '0 0 2 * * ?';
System.schedule('Account Rating Nightly', cron, new AccountRatingScheduler());
// every Sunday at 1:30 AM
System.schedule('Account Rating Weekly', '0 30 1 ? * SUN', new AccountRatingScheduler());

Salesforce allows a maximum of 100 scheduled Apex jobs in an org at one time. You can also schedule from the user interface under Setup, Apex Classes, Schedule Apex, but that path only offers daily and weekly options, so anything more precise needs System.schedule.

A class that implements both Database.Batchable and Schedulable can be scheduled directly without a separate scheduler class. Keeping them separate is usually cleaner, because the scheduler then controls the scope value without touching the batch.

Testing a batch class

Two rules govern batch tests. The job must run between Test.startTest() and Test.stopTest(), and inside a test only a single execute runs.

That second rule decides how much test data you create. If you insert 300 records and run with a scope of 200, only the first 200 are processed and your assertion on all 300 fails. Keep the test data at or below the scope.

@isTest
private class AccountRatingBatchTest {

    @testSetup
    static void makeData() {
        List<Account> accounts = new List<Account>();
        for (Integer i = 0; i < 50; i++) {
            accounts.add(new Account(
                Name = 'Batch Test ' + i,
                AnnualRevenue = 20000000,
                Rating = 'Cold'
            ));
        }
        insert accounts;
    }

    @isTest
    static void highRevenueAccountsBecomeHot() {

        Test.startTest();
        Database.executeBatch(new AccountRatingBatch(), 200);
        Test.stopTest();   // the batch runs to completion here

        Integer hotCount = [SELECT COUNT() FROM Account WHERE Rating = 'Hot'];
        System.assertEquals(50, hotCount, 'All 50 accounts should be rated Hot');
    }

    @isTest
    static void lowRevenueAccountsBecomeWarm() {

        List<Account> accounts = [SELECT Id FROM Account];
        for (Account acc : accounts) {
            acc.AnnualRevenue = 500000;
        }
        update accounts;

        Test.startTest();
        Database.executeBatch(new AccountRatingBatch(), 200);
        Test.stopTest();

        Integer warmCount = [SELECT COUNT() FROM Account WHERE Rating = 'Warm'];
        System.assertEquals(50, warmCount, 'All 50 accounts should be rated Warm');
    }
}

Test.startTest() gives the test method a fresh set of governor limits, so your setup data does not eat into what the batch is allowed to use. Test.stopTest() forces the asynchronous job to run synchronously and finish before the next line executes, which is why your assertions can sit immediately after it. Assertions placed before Test.stopTest() will fail, because the batch has not run yet.

A batch class needs both paths of its logic covered, not just the happy one. Write a test for a record that fails the update as well, so the error branch is exercised.

Reading AsyncApexJob to check status

Every asynchronous job writes a row to AsyncApexJob. This is your first stop when someone says the nightly job did not run.

List<AsyncApexJob> jobs = [
    SELECT Id, ApexClass.Name, Status, ExtendedStatus,
           JobItemsProcessed, TotalJobItems, NumberOfErrors,
           CreatedDate, CompletedDate, CreatedBy.Name
    FROM AsyncApexJob
    WHERE JobType = 'BatchApex'
    ORDER BY CreatedDate DESC
    LIMIT 20
];

for (AsyncApexJob job : jobs) {
    System.debug(job.ApexClass.Name + ' | ' + job.Status +
                 ' | ' + job.JobItemsProcessed + ' of ' + job.TotalJobItems +
                 ' | errors: ' + job.NumberOfErrors +
                 ' | ' + job.ExtendedStatus);
}

What the fields tell you:

  • Status moves through Holding, Queued, Preparing, Processing, and then Completed, Failed or Aborted.
  • TotalJobItems is the number of chunks. JobItemsProcessed is how many have finished, which gives you progress.
  • NumberOfErrors counts chunks that threw an unhandled exception.
  • ExtendedStatus carries the first error message, and it is usually the single most useful field on the record.

A status of Completed with a non-zero NumberOfErrors is the case people miss. The job finished, nobody was told, and a quarter of the records were never updated.

The same information is visible in Setup under Apex Jobs. Knowing the object name and being able to query it is what an interviewer is listening for, because the trainer observation holds here too: candidates cannot debug. "Your batch shows Completed but the records did not change, what do you check?" tells the panel far more than a definition of the Batchable interface.

Six things that break batches in production

1. A non-selective query in start. On an object with millions of rows, a WHERE clause on an unindexed field either times out or throws a non-selective query exception. The fix is to filter on an indexed field, which means Id, Name, OwnerId, CreatedDate, LastModifiedDate, RecordTypeId, a Master-Detail or Lookup field, an External Id, or a Unique field. If no such field fits, ask Salesforce Support for a custom index, or narrow the job to a date window and run it repeatedly.

2. Scope too large for the work in execute. The job fails partway with an Apex CPU time limit exceeded or an Apex heap size too large error. The fix is to lower the scope. Start at 50 and work down. If the object carries heavy triggers and flows, your execute is only part of what runs in that transaction.

3. Assuming static variables carry across chunks. A counter in a static variable reads as the per-chunk count, not the job total, and the summary email reports nonsense. The fix is Database.Stateful with instance variables, and keeping only aggregates in them so serialisation does not blow the heap later in the run.

4. An unhandled exception in execute killing only that chunk. Batch Apex does not stop when a chunk throws. The chunk's DML is rolled back, NumberOfErrors goes up, and the job carries on to the next chunk and reports Completed. The fix is a try/catch around the per-record work, Database.update(list, false) instead of plain DML, and writing the failed Ids to a custom logging object so there is something to re-run.

5. Hitting the five-concurrent-jobs limit. Only five batch jobs can be queued or active at one time. Beyond that they sit in the Apex flex queue in Holding status, up to 100, and a direct Database.executeBatch call can throw a LimitException. This shows up when several schedulers are all set to 2:00 AM. Stagger the cron times, or chain the jobs from finish so they run in sequence.

6. Triggers and flows on the target object consuming the chunk's limits. Your execute might be clean, but the update fires a trigger that queries related records, which fires a flow, which fires another trigger. All of it runs inside your chunk's 60 seconds of CPU and 150 DML statements. The fix is to skip records that do not need changing, use a static bypass flag or a custom setting so automation is suppressed during bulk jobs, and lower the scope so the combined work fits.

What the interviewer is actually checking

  • Whether you understand that each execute is a separate transaction with fresh limits, and can say what that buys you. This is the concept the whole feature rests on.
  • Whether you know why QueryLocator allows 50 million records and Iterable stops at 50,000, because that single fact decides whether a one million record job works at all.
  • Whether you reach for the scope parameter when a job fails on CPU or heap, and whether you know it does not help a non-selective query.
  • Whether you can name what Database.Stateful preserves and what it does not. Saying "instance variables yes, statics no" answers it in one line.
  • Whether you can debug a batch that reports Completed but did nothing. Memorised definitions collapse the moment the question turns into a real project scenario, and this is the question that does it.

Where to take this next

Batch Apex is one piece of the asynchronous toolkit, and the same interview usually moves straight on to Queueable, Schedulable, future methods and Platform Events. The way to be ready is to build a batch job against real volume in a developer org, break it deliberately by raising the scope until it fails, then read the AsyncApexJob row and explain what happened. That story, told with the object names and numbers in it, is what makes an answer credible.

If you want that practice with a trainer looking over the code and asking the follow-up questions a panel would ask, the Salesforce training in Pune with placement support programme at RizeX Labs covers asynchronous Apex as a hands-on module rather than a slide.

Last updated 4 October 2026