Thursday, 7 November 2013

Singly Linked Lists and Memory Alignment - Access Violation Exception

I came across a interesting case, whereby a singly linked list wasn't correctly aligned for the MEMORY_ALLOCATION_ALIGNMENT boundary on x64 systems. My analysis and explanation of why it may have happened is within the thread, which will be posted below with some MSDN API documentation for you to study.

BSOD 0x0000001E, 1D, 3B on Computer Shutdown

Additional Information:

InterlockedFlushSList 

_aligned_malloc

How to use Interlocked Singly Linked Lists?

I've also opened a thread at Sysnative, regarding if it is possible to XOR a Singly Linked List, since I noticed a XOR Assembly instruction in one of the instructions.

[Question] Possible XOR a Singly Linked List?




Wednesday, 6 November 2013

Fast Mutexes, Guarded Mutexes and Semaphores

The next blog post of my exploration of kernel-mode synchronization mechanisms (I'm not going to bother with user-mode mechanisms), in this blog post, I'm going to talk about the Mutex object. Mutexes are a object which provide a form of mutual exclusion, they are very similar to Semaphores. 

Mutexes only allow one thread which holds the mutex object, access to a certain protected resource.


Source - Basics of Mutexes and Spinlocks


Fast Mutexes:

Fast Mutexes (or Executive Mutexes) are built a dispatcher object called a Event object, a Event object can be within two states: Signaled or Non-Signaled, this thereby applies to a Fast Mutex. A Mutex comes Signaled, when it is not owned by a thread, and therefore can be a acquired. A Mutex is said to be Non-Signaled, when the object is owned by a thread, when the object is released, the Mutex object changes to Signaled, and the thread is released.

Fast Mutexes disable normal Kernel-Mode APCs when raising the IRQL Level of the processor to IRQL Level 1 or APC_LEVEL. Fast Mutexes can't be acquired recursively, that is to say, the same thread can acquire the same Mutex multiple times, but the thread must release the Mutex the same number of times it acquired it, otherwise the system will run into a deadlock.

A reetrant mutex is recursive.

Guarded Mutexes:

Guarded Mutexes are similar to Fast Mutexes, but use a different kind of dispatcher object, this object is called a Gate object. Any thread which holds a Guarded Mutex, runs inside a Guarded Region, which in turn disables all APC activity (Special and Normal).

When a Gate object, is set to Signaled, when the thread signals the gate; one waiting thread is then released.

More Information - Fast Mutexes and Guarded Mutexes

Semaphores:

Unlike, Mutex objects which generally only allow one thread to have access to a resource, Semaphores can allow multiple threads at the same time to have access to that resource. A semaphore keeps a count of how many threads are accessing that resource, and how many threads are allowed to have access to a that resource. 

The Semaphore can be thought of, as a gate, in that it controls the number of threads to a resource. A Semaphore can be Signaled or Non-Signaled, when Signaled the thread releases it's access to the resource and the Semaphore count is incremented, meaning that a Semaphore object has become free to use and allow access to the protected resource. 

When the Semaphore is obtained by another thread, the state changes from Signaled to Non-Signaled, and the Semaphore count is decremented. Once the count has become zero, no more threads are able to wait for the thread to become signaled.

More Information - Semaphore Objects

Associated Problems:

Deadlocks: One thread obtains a resource , and then another thread obtains a resource, although they are both waiting to acquire each others resources. One thread can continue and leave it's wait state, until it's obtained the other resource which is being held by the other thread.


Priority Inversion: A thread with a higher-priority is waiting for a lower-priority thread; this breaks the mechanism of thread priorities. 

Resource Starvation: A process is never given the resources it requires to complete any tasks assigned to it, leading to the process freezing.

References: 

Mutex Vs Semaphore

Muxtexes Vs Semaphores - Semaphores: Part 1

Difference between binary Mutex and Semaphore

 





Spinlocks and Queued Spinlocks

You will notice very often, there is always some form of synchronization mechanism within each call stack you examine from a dump file. The next few blog posts, will attempt to explain these synchronization mechanisms, and potentially how they can cause bugchecks if not used correctly. In this blog post I'm going to explain spinlocks and queued spinlocks.

Spinlocks:

Spinlocks are a form of locking mechanism which is used by the kernel to enforce the concept of mutual exclusion. Typically, spinlocks are used in conjunction with critical regions, and are used to allow only one thread to execute code within that critical region. Spinlocks are used at IRQL Level 2 or Dispatch Level, and therefore are allocated with non-paged pool.

Spinlocks are given the associated level of IRQL Level 2, in order to mask or preempt any dispatching mechanisms and interrupts at that level or below, to allow the thread within the critical region to execute the code much more quickly, this is because when another thread is attempting to acquire the spinlock, the thread and the processor (where the thread is running) is kept waiting (or spinning) until the lock is obtained. This is a expensive operation, and is usually shown as 100% CPU usage.

Spinlocks are generally acquired by using a test and set operation (written in Assembly) which is completed within one atomic instruction, to prevent other threads from interrupting and then acquiring the spinlock. The lock bts Assembly instruction is typically used, to lock the processor bus and stop other processors from interrupting. The test and set operation is used to test the lock variable, here is a example of a spinlock acquisition in Assembly code:

Source: Spinlock Wiki Article


 Since, using Spinlocks can be expensive, a Intel based Assembly instruction can be introduced, called pause. This is designed to reduce power consumption and prevent the CPU from having to re-order many read requests from waiting threads when the Spinlock is released.

It's important to remember, that any thread holding a Spinlock, isn't able to cause any page faults or call any dispatch routines, since this will result in a system crash and most likely bugcheck.

Queued Spinlocks:

As you may have established, Spinlocks are not the most efficient locking mechanisms in terms of performance, and as a result queued spinlocks were created exclusively for the Kernel. They are not supported for the use by third-party driver developers.

A queued spinlock works similarly to a conventional spinlock, however, when a processor and it's thread wish to acquire a spinlock being held, they are  placed in a queue for that spinlock. The queue uses a FIFO (First In - First Out) ordering mechanism, the processor with the spinlock will give the spinlock to the next processor identifier in the queue. The processor technically adds it's processor identifier to the queue. 

We can view the number of global queued spinlocks with the !qlocks extension:


There are currently no acquired queued spinlocks in this example, however, the owner will be specified if there is such a spinlock.

Instack Queued Spinlocks:

These are the queued spinlocks which have been made to third-party driver developers, they still use the same _KSPIN_LOCK data structure, but these spinlocks use a handle to a data structure called KLOCK_QUEUE_HANDLE, the handle is local to the stack of the calling thread.

Spinlock Common Problems:

A list of problems and their preventions with use of Spinlocks can be found here on the MSDN page - Preventing Errors and Deadlocks While Using Spinlocks 

More References:

Synchronicity - A Review of Synchronization Primitives
Tools of the Trade - A Catalog of Synchronization Mechanisms  


















Tuesday, 5 November 2013

IDT Hooks

I came across a interesting question on the KernelMode forum, with regards to IDT hooking and if it's malicious within the context in which they posted the question.

I've decided to investigate further into the concept of IDT hooking, and why some software developers may wish to use this concept, even though it isn't advised by Microsoft, who have in fact protected the kernel from IDT hooking with Kernel Patch Protection (x64 systems only). The KPP was added, since IDT hooking can cause system instability, since this kind of activity isn't supported by the operating system. Any attempt to hook into IDT on x64 systems will lead to a Stop 0x109 bugcheck.

KPP doesn't protect against hooking of IDT or System Service table, since some software is dependent upon those features.

Some developers may wish to change the behavior provided by the SSDT (System Service Dispatch Table).

Hooking the IDT:

Hooking the IDT: Programming
Hooking Software Interrupts

If you have read through the two above articles, you may be wondering what is the IDTR register used for, and what are interrupt dispatch gates?

The IDTR register is simply used to store the IDT table, and help to locate the memory address in which the appropriate vector is residing. Vectors are our exception codes and interrupt numbers within the IDT. The interrupt dispatch gates are essentially the same thing as our interrupt descriptors within the IDT, they can be a interrupt gate (hardware interrupts), trap gates (software interrupts) or task gates (switches TSS and enables different thread/process to have control of the processor).

Malicious uses of IDT and SSDT Hooking:

Hooking of the INT2E handler for generic System calls, which allows User-Mode threads to transition into Kernel-Mode threads, could allow rootkits to have single hooking point on System calls. 

Hooking of the SSDT, could enable rootkits to have control over the I/O of calls from User-Mode, and allow modification of certain processes and registry keys.



 







Sunday, 3 November 2013

Object Retention - Object Manager

Objects can be temporary or permanent, retention of permanent objects is quite simple, they are not deleted. Temporary objects have two phrases of retention. We should understand that when a process acquires a object, the reference count (handle count + pointer count) is incremented by 1, and when that handle is closed, then the reference count is decremented by 1.

When,  the handle count of an object drops to 0, the Object Manager removes the object's name from the global namespace, therefore stopping any new processes from opening handles to that object.

Once, the name has been removed, then the object will be only deleted, once the reference count has dropped zero, since kernel processes are able to use object with pointers, hence the reason why there is a pointer count field within the object header data structure.

The reference count is a combination of the pointer reference count and handle reference count. 

We can use the !object extension to view the above mentioned fields.


 The reference count would be 48.

It's important to remember that objects which are using paged pool, must only be freed when the IRQL Level is below 2, since page faults will be illegal operations, and thus will cause the system to crash.


 

Object Security - Object Manager

As promised, I'm going to explain some other topics relating to Objects and the Object Manager.

When a process is created, and then wishing to use the object, either through opening a handle or using a pointer (reserved for kernel), then it must inform the Object Manager of which access rights it wishes to acquire. For example, opening a handle to a file object (can be a storage device), it may wish to read or write to that device. If so, the Object Manager will need to call the Security Reference Monitor, and show the desired access rights of the process, if the object's security descriptor permits these access rights, then the process gains a granted access rights.

We can view the access rights of a object using WinObj, ensure your running the program as a Administrator, by clicking File and and then Run As Administrator.



Here is an example of using a Object directory, and then viewing it's security access rights. Right-Click the Object Type/Directory, and then select Properties.


We can view the Security Descriptor of a object using WinDbg, I'm using a process object in this example:

Enter the !process extension, to obtain the address of the process object, to be used with the !object extension.


Now, use the !object extension with the address of the process object, this will give us the address of the Object Header data structure, which will contain the Security Descriptor field, which in turn will contain the address of the Security Descriptor to be used with the !sd extension.


Use !sd extension with the address in the Security Descriptor field to obtain all the Security Descriptor information. Since, the !sd extension didn't work in my Kernel Memory dump, I've taken the example from the WinDbg documentation.



I'm unlikely to be able to explain every detail of how Security Descriptors are formed, and all the internals of Object Security, since it's a wide topic for a blog post.

I'll explain some of the fields for the !sd extension:

Revision: Version of the SRM (Security Reference Monitor) security model. 

Flags: Characteristics of the security descriptor.

Owner: Owner SID, or security ID.

Group: Group SID for primary group for the object.

The flags which are currently set are:

SE_DACL_PRESENT: Indicates that the security descriptor has a discretionary access control list present (DACL), which shows who has access to that object. If this flag is not set or NULL, then everyone has full access to that object.

SE_SELF_RELATIVE: Indicates that the security descriptor has all the security information in one continuous block of memory (most likely a array).

We can also view security descriptor information in Process Explorer.













Debugging Stop 0x1 - APC's and Guarded Regions

This is the first Stop 0x1, which I actually came across on my own, and to be honest is one of those bugchecks which doesn't contain any information at all really, as a result of when the bugcheck is produced. To really understand, how this bugcheck works and why it happens, you need to understand the concept of APCs and their types, and also how they are used with Critical Regions and Guarded Regions.

APCs are quite a lengthy subject, and therefore I will not explain completely how they work and their internals, but will provide some useful references for you to read or take note of. They are wonderfully explained in the Windows Internals book, which I absolutely recommend that you purchase.

Brief Explanation of APCs  

APC's are a form of asynchronous interrupt, and run within the context of a particular thread or process address space. They can allow page faults, call system services, wait for objects and their handles and acquire objects. 

APCs are called APC objects, and when a thread wishes to use a APC, a APC object is inserted into the queue of that thread, called a APC Queue. The APCs are then executed when the IRQL Level is 1.

There are two main types of APCs: Kernel-Mode and User-Mode. These two types are then divided into Normal and Special. Kernel-Mode APCs can simply run in the context of the target thread without having to wait to gain permission from that thread, whereas, User-Mode APCs have to wait for permission.


APCs and Stop 0x1

Getting back to the subject, of the relationship between APCs and Stop 0x1, let's start examining the important points within the dump file.


This bugcheck always occurring exiting a Service Call, which in this case is usually from calling KeExitGuardedRegion.

 As already pointed out, in the description of the bugcheck, the most significant parameter is the 3rd parameter, we indicates the current value of the thread's CombinedAPCDisable field. The parameter is split into two 16-bit values, a SpecialAPCDisable value and a KernelAPCDisable value.

We can clearly see that both values are negative, which therefore shows that both Special APCs and Kernel APCs were disabled but never re-enabled again. Since both APC types have been disabled, the thread would have entered a Guarded Region rather than a Critical Region, since no APCs are executed within that thread's context upon entering a Guarded Region.

Device drivers will enter Guarded Regions and Critical Regions (disabling APCs), usually when holding a lock (there are many different types of locks), to prevent any Kernel-Mode APCs being used to suspend or terminate the thread, if the thread was terminated or placed into a wait state then the system could potentially deadlock and hang.

When APCs are disabled, two/three fields are set in the _KTHREAD data structure as shown below:



When exiting a Guarded Region or a Critical Region, APCs must be re-enabled again, since they are largely used I/O Manager and in I/O operations. 

Another few points, the second parameter contains the value of the thread's APCStateIndex field, which is stored in the _KAPC data structure:

The APCStateIndex field is a pointer to the APCState field found in the _KTHREAD structure:

We can clearly see that the APCState field contains another data structure called _KAPC_STATE. SavedAPCState which is part the _KTHREAD has the same data structure stored.




The APCState field is called a APC Environment, and this field is used for APCs targeted at the current thread's context, and does not take in regard, if the thread is running within it's own process or attached to another process. SavedAPCState field is also a APC Environment, and contains APCs at threads which are not running within the context of a current process, and therefore these APCs must wait to be delivered.

Driver Verifier and Critical Region Counts

The best option you have, is running Driver Verifier and then checking the Critical Region log file. By using the !verifier 0x200 flag, and then finding any mismatched calls. Critical Regions counts are explained in the link in this paragraph.

Overall, your best option would be to run Driver Verifier.