Showing posts with label cpp_link. Show all posts
Showing posts with label cpp_link. Show all posts

Dec 1, 2024

[C++][cpponsea] Keep C++ Binaries Small - summary

Reference:
https://youtu.be/7QNtiH5wTAs?si=c8I_prA9k9Hasfdr


Problem

  • Devices with limited storage
  • Devices with limited bandwith
  • Env impact

Parts of an executable

  • Header
  • Text
  • Data (init. vs uninit)
  • Read-only data
  • symbol table
  • Relocation info
  • (ref https://eli.thegreenplace.net/2013/07/09/library-order-in-static-linking)
  • Import table / Procedure linkage table
  • Export table
  • Exception handling info
  • Debugging info

Consider

  • cross platform?
  • linking/symbol resolution handling
  • PIC
  • Endianness
  • Architecture / extensibility

Tools


Coding part


Object init.

0(.zerofill) vs. other thing else; decrease size


Member ordering

Usually won't reduce the binary size due to compiler is able to optimize.


Special member functions

  • inline default special functions
  • inline empty special functions
  • following the rule of zero(RoZ)
  • with either virtual or non-virtual destructor


templates

  • minimal template code section; same as marco
  • Exact/promote to the upper level, which the code is not generated
    for every new type instance.

Compiler options

  • -Os
  • -flto
  • -fdata-sections / -ffunction-sections -Wl, --gc-sections
    • The Traditional Way (Still Valid) If you are linking against a traditional static library (.a file), the linker still operates on an object-file level. If a static library is built with strlen.o, strcpy.o, and printf.o as separate files inside the archive, and your code only references strlen, the linker will strictly pull in strlen.o. The other object files are completely ignored, keeping your binary small. Many embedded and legacy C libraries are still structured exactly this way.
    • The Modern Way: Section Garbage Collection Today, instead of splitting source code into a million files, we let the compiler and linker do the heavy lifting using function-level sections. When compiling modern C code (using GCC or Clang), you can pass specific flags: -ffunction-sections and -fdata-sections: This tells the compiler to put every single function and data item into its own distinct section inside a single object file, rather than bundling them all into one giant .text section. --gc-sections (Linker flag): This tells the linker to perform "garbage collection" and throw away any unused sections during the final build. Why this matters: You can have a single string.c file containing 50 functions. With these flags, the linker will still grab only strlen and discard the other 49, achieving the exact same tiny binary size without the architectural headache.
    • Link-Time Optimization (LTO) The ultimate evolution of this is LTO (compiled with -flto). With LTO enabled, the compiler doesn't just emit machine code into object files; it emits its internal intermediate representation. At the linking stage, the compiler looks at the entire program holistically. It can inline strlen directly into your code, eliminate dead code with extreme precision, and completely optimize away unused functions, regardless of how the source files or libraries were structured.
  • -Wl, --icf=all or -Wl, --icf==safe
  • -Wl, -s and -Wl, --strip-all
  • -Wl, --as-needed
  • -mllvm -inline-threshold=<n>
  • -fstack-protector
    https://clang.llvm.org/docs/ClangCommandLineReference.html#cmdoption-clang-fstack-protector
  • -finline-limit=n
    https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html#index-finline-limit
  • Identical Code Folding (ICF)
    https://research.google/pubs/safe-icf-pointer-safe-and-unwinding-aware-identical-code-folding-in-gold

Nov 21, 2024

[C++] key function

Reference:
undefined reference to vtable for X
What is a C++ "Key Function" as described by gold?



Every polymorphic class requires a virtual table or vtable describing the type and its virtual functions. Whenever possible the vtable will only be generated and output in a single object file rather than generating the vtable in every object file which refers to the class (which might be hundreds of objects that include a header but don't actually use the class.)

The cross-platform C++ ABI states that the vtable will be output in the object file that contains the definition of the key function, which is defined as the first non-inline, non-pure, virtual function declared in the class.

If you do not provide a definition for the key function (or fail to link to the file providing the definition) then the linker will not be able to find the class' vtable.


class X {
public:
  virtual ~X() = 0;
  void f();
  virtual void g() { }
  virtual void h(); // defined inline out side of class.
  virtual void i(); // <-- Key fuction.
  virtual void j();
};

inline void X::h() { }

reproduce the error:

#include <iostream>

class Base {
public:
    virtual void foo() = 0; // Pure virtual function
};

class Derived : public Base {
public:
    void foo() override; // Declaring but not defining foo()
};

int main() {
    Derived d; // This will cause a linker error due to missing vtable
    return 0;
}

Sep 17, 2022

[elf][note] entry function

ELF executable entry function

_start function facts


_start function implements

  1. Early low-level initialization, such as
    1. Configuring processor registers
    2. Initializing external memory
    3. Enabling caches
    4. Configuring the MMU
  2. Stack initialization, making sure that the stack is properly aligned per the ABI requirements
  3. Frame pointer initialization
  4. Initialization of the C/C++ runtime
    Relocate any relocatable sections (if not handled by the loader or linker)
    1. Initializing global and static memory
    2. Runtime initializes a subset of uninitialized memory (no = in the declaration) to 0.
      This includes global and static variables, but not stack variables. All uninitialized data that needs to be set to 0 is placed into the .bss section of the compiled program image by the linker. The location of the .bss section is identified during initialization, and the memory is typically set to 0 with memset.
    3. C++ global objects must be constructed before calling main. The linker places these constructors into the .init, .init_array, or .ctors section of the image.
      Some compilers also allow C and C++ functions to be marked as a constructor using a compiler attribute (e.g., __attribute__((constuctor))). The constructors are stored in a list by the linker.
      The runtime initialization process iterates through the list and calls each constructor.
    4. Prepare the argc and argv variables for invoking main (even if it’s just setting these to 0/NULL)
    5. Perform any additional setup steps required by the C/C++ standard library implementation.
      These additional runtime initialization steps are run for many programs (but not all):
      1. Heap initialization
      2. Initialize stdio (i.e., stdin, stdout, stderr)
      3. Initialize exception support (if using C++)
      4. Register destructors and other cleanup functions that will run when exiting the program (using atexit and __cxa_atexit)
      5. Assembly files commonly found during this portion of the startup process are crtbegin.s, crtend.s, crti.s, and crtn.s.
      6. Prepare environment variables
  5. Initialization of other scaffolding required by the system
      Program scaffolding setup before main might include:
      1. Threading support and thread local storage
      2. Buffer overrun detection
      3. Stack logging
      4. Run-time error checks
      5. Locale settings
      6. Math error handling
      7. Default math library precision
  6. Jumping to main
  7. Exiting the program with the return code from main

So, how do we get to the _start?

  • Baremetal: Reset Vector
  • Bootloader Launches Application
  • OS Calls an exec function
    Loaders will often perform the following actions:
    • Check permissions
    • Allocate space for the program’s stack
    • Allocate space for the program’s heap
    • Initialize registers (e.g., stack pointer)
    • Push argc, argv, and envp onto the program stack
    • Map virtual address spaces
    • Dynamic linking
    • Relocations
    • Call pre-initialization functions

Jul 11, 2022

[Linkage] COMDAT

Reference:
https://maskray.me/blog/2021-07-25-comdat-and-section-group


COMDAT gives indication to the linker to deduplicate the inline/weak reference/external linkage symbols.


https://github.com/llvm/llvm-project/blob/2e603c6/clang/lib/AST/ItaniumMangle.cpp#L1458-L1472

inline vs. non-inline name mangling with extra letter 'L'.

e.g.

_ZSt8in_place vs. _ZStL8in_place

Jan 27, 2019

[split stack] reading notes and references

Reference:
gccgo split stack implementation
  1. The stack can start splitting at any point.
  2. The stack size is automatically recorded at program startup,
    and each thread startup.
  3. The gold linker detects calls from split-stack code to non-split-stack
    code, and rewrites the function header to force a large stack segment to be allocated.
    i.e.
    When not using the gold linker, calls from split-stack code to non-split-stack code will just have whatever is left of the current stack segment, which may not be large enough.
    (look up to "Backward compatibility" section)


In the complex GCC ecosystem the linker is separate from the compiler.
GCC can't assume that gold is available at all.
When building gccgo, configure using
--with-ld=/path/to/gold

The -fuse-ld=gold option is newer than gccgo.
Ian supposes it would be nice if:
* the GCC configure process checks whether -fuse-ld=gold works; if so:
  * -fuse-ld=gold is passed to the libgo configure/build
  * -fuse-ld=gold is used by default by the gccgo driver program



Reference:
Split Stacks in GCC




Obvious benefits

  • The memory usage of a typical multi-threaded program can decrease significantly, as each thread does not require a worst-case stack size.
  • It becomes possible to run millions of threads
    (either full NPTL threads or co-routines) in a 32-bit address space.




Basic explained

Stack will have a guaranteed zone which is always available.
Reference:
[LWN] Preventing stack guard-page hopping


The size of the guard area will be target specific.
It will include enough stack space to actually allocate more stack space.
Each function will have to verify that it has enough space in the current stack to execute.

The basic verification will be a comparison between the stack pointer and the current bottom of the stack plus the guaranteed zone size.
This will have to be the first operation in the function, and will also be target specific.

It must be fast, as it will be executed by each called function.

Two cases to consider.
  1. For functions which require a stack frame less than the size of the guaranteed guard area, we can do a simple comparison between the stack pointer and the stack limit.
  2. For functions which require a larger stack frame, we must do a comparison including the size of the stack frame.




Design options

  1. Reserve a register to hold the bottom of the stack plus the guaranteed size. This will have to be a callee-saved register.
  2. Use a TLS(Thread Local Storage) variable. In the general case, in a shared library, this will require calling the __tls_get_addr function.
    Reference:
    How fast is thread local variable access on Linux
    (GOLD elf linker)
    http://gittup.org/cgi-bin/man/man2html?gold+1 

    That means that that function will have to work without requiring any additional stack space.
    This is infeasible unless the whole system is compiled with split stacks.
    It would require dlopen's LD_BIND_NOW to be set, so that the __tls_get_addr function is resolved at program startup time.
    Even that is probably insufficient unless we can ensure that the space for the (TLS) variable is fully allocated.
    In general Ian doesn't think they can ensure this, because dlopen can cause a thread to require more space for TLS variables, and that space will be allocated on the first call to __tls_get_addr.
    Reference:
    http://man7.org/linux/man-pages/man8/ld.so.8.html
    LD_BIND_NOW (since glibc 2.1.1)
    If set to a nonempty string, causes the dynamic linker to
    resolve all symbols at program startup instead of deferring
    function call resolution to the point when they are first
    referenced.  This is useful when using a debugger.
  3. Have the stack always end at a N-bit boundary.
    E.g., if we always allocate stack segments as a multiple of 4K,
    then align each one so that the stack always ends at a 12-bit boundary.
    Then the amount of space remaining on the stack is SP & 0xfff.
  4. Introduce a new function call which handles the comparison of the stack pointer and the stack expansion.
  5. Reuse the stack protector support field.
    When using glibc each thread descriptor has a field used by the stack protector.
    Of course it is then not possible to use split stacks in conjunction with stack protector.
  6. At least on x86, arrange to allocate a new field in the TCB(thread control block) header accessible via %fs or %gs.
    This is probably the best solution, and it is the one implemented for i386 and x86_64.

Reference:
TCB Thread Control Block in linux kernel:
https://en.wikipedia.org/wiki/Thread_control_block



Expanding the stack

  • Expanding the stack requires allocating additional memory.
  • This additional memory will have to be allocated using only the stack space slot.
  • All of the functions used to allocate additional stack space must be compiled to not use a split stack.
  • A new function attribute, no_split_stack will be introduced to mean that the stack should not be split.
  • It would also work to ensure that the stack is large enough that they do not need to split the stack during the allocation call.
  • After expanding the stack, the function will copy any stack based parameters from the old stack to the new stack.
  • Fortunately, all C++ objects which require a copy or move constructor are implicitly passed by reference,so copying the parameters on the stack is OK.
  • For varargs functions, this is impossible in general, so we will compile varargs functions differently:
    they will use an argument pointer which is not necessarily based on the frame pointer.
    For functions which return objects on the stack, the objects will be returned on the old stack. (RVO)
    This should normally happen automatically, as the initial hidden parameter will naturally point to the old stack.
  • When expanding the stack, the return address of the function will be managed to point to a function which will release the allocated stack block and reset the stack pointer to the caller.
    Reference:
    http://vsdmars.blogspot.com/2017/11/assembly-note.html
  • The address of the old stack block, and the old stack pointer, will have been saved somewhere in the new stack block.




Backward compatibility

We want to be able to use split stack programs on systems with pre-built libraries compiled without split stacks.
This means that we need to ensure that there is sufficient stack space before calling any such function.

Each object file compiled in split stack mode will be annotated to indicate that the functions use split stacks.

This should probably be annotated with a note but there is no general support for creating arbitrary notes in GNU as.

Therefore, each object file compiled in split stack mode will have an empty section with a special name: .note.GNU-split-stack

If an object file compiled in split stack mode includes some functions with the no_split_stack attribute, then the object file will also have a .note.GNU-no-split-stack section.

This will tell the linker that some functions may not have the expected split stack prologue.

When the linker links an executable or shared library, it will look for calls from split-stack code to non-split-stack code.

This will include calls to non-split-stack shared libraries
(thus, a program linked against a split-stack shared library may fail if at runtime the dynamic linker finds a non-split-stack shared library;
it might be desirable to use a new segment type to detect this situation).

For calls from split-stack code to non-split-stack code, the linker will change the initial instructions in the split-stack (caller) function.
This means that the linker will have to have special knowledge of the instructions that the compiler emits.
The effect of the changes will be to increase the required frame-size by a number large enough to reasonably work for a non-split-stack.
This will be a target dependent number; the default will be something like 64K.
Note that this large stack will be released when the split-stack function returns.
Note that I'm disregarding the case of split-stack code in a shared library calling non-split-stack code in the main executable; that seems like an unlikely problem.


Function pointers are a tricky case.
In general we don't know whether a function pointer points to split-stack code.
Therefore, all calls through a function pointer will be modified to call (or jump to) a special function __fnptr_morestack.
This will use a target specific function calling sequence, and will be implemented as though it were itself a function call instruction.
That is, all the parameters will be set up, and then the code will jump to __fnptr_morestack.
The __fnptr_morestack function takes two parameters: the function pointer to call, and the number of bytes of arguments pushed on the stack.

Jan 1, 2019

[linker] Linker Namespaces (glibc >= 2.3.4)

[Go] Execution modes

Reference:

Go Execution Modes - Ian Lance Taylor  https://goo.gl/mrzwCz
https://golang.org/cmd/link/
https://golang.org/cmd/go/#hdr-Build_modes
https://golang.org/cmd/go/#hdr-Compile_packages_and_dependencies
https://github.com/golang/go/issues/18246
https://blog.ksub.org/bytes/2017/02/12/exploring-shared-objects-in-go/


Legacy Go(v1.4.0) support 3 execution modes:

  1. A statically linked Go binary.
    Default for program does not import the net or os/user packages and does not use cgo or SWIG.
  2. A dynamically linked Go binary.
    Default for program imports the net or os/user packages and does not otherwise use cgo or SWIG
  3. A Go binary linked with arbitrary non-Go code.
    Default for program uses cgo or SWIG.
    The interface between Go and non-Go code is a C style API
    (SWIG permits calling between C++ and Go, but this is implemented using a C style API).
    Can be selected with -ldflags -linkmode=external. (obsolete)


API styles

  1. Go
  2. C, i.e extern "C"



Design in mind

  • Go code that is combined into a single executable image must be built with the same version of the Go toolchain.
  • We require further that if any Go package appears more than once in the executable image,
    it must be built from the same source code. (Same as C++'s header file)
  • This restriction comes from both ways: Golang <-> C



Go runtime:

  • All Go code shares a single runtime. 
  • All Go code uses the same memory allocator.
  • The same goroutine scheduler.
  • In general acts as though it were linked into a single Go program.



New execution modes:

  • Go code linked into, and called from, a non-Go program. (Go v.1.5.0)
    Go code acts as a library that may be called by a non-Go program.
    A single Go library will be an archive, a .a file on Unix,
    or as a shared library, a .so file on Unix, providing a C style API.
    This mode supports people who must work with large existing programs, especially in C/C++.
    It permits them to extend those existing programs with new packages written in Go.
    i.e the binary code is sheer an ELF format.
  • Go code linked into a shared library loaded as a plugin by a program (Go or non-Go) that supports a C style plugin API. (Go v1.6.0)
    i.e dlopen/dlmopen (dlmopen appears in glibc 2.3.4 for linker's namespace)
    https://sourceware.org/glibc/wiki/LinkerNamespaces
    Golang binary code in .so ELF format can be dlopened by C/C++.
  • Go code linked into a shared library loaded as a plugin by a Go program that supports a general Go style plugin API.
    A shared library can be dlopened by Golang.
    https://golang.org/pkg/plugin/
  • A Go program that uses a plugin interface, either C style or Go style, where plugins are implemented as shared libraries.
    Golang can dlopen any C/C++ .so libraries.
  • Building a Go package, or collection of packages, as a shared library that may be linked into a Go program. (Go v1.6.0)
    Single/Multiple Golang packages build into single shared library, which can be linked to other Golang program.
    Updating the Go run-time to a new version requires rebuilding all Go programs that use it.
  • A Go program built as a PIE--a Position Independent Executable. (Go v1.6.0)
    Go program is built as usual, but the resulting executable is position-independent, and may be relocated at run time.
    (i.e -fPIC in C/C++)



Go tool flags:

  • -buildmode
archive:
Default build mode for a package that is not main.
Builds the package into a .a file.

c-archive: (Go v1.5.0)
Requires a main package, but the main function is ignored (init functions are run as usual).
Build the main package, plus all packages that it imports, into a single C archive file.
The only callable symbols will be those functions marked as exported.

shared: (Go v1.6.0)
Combine all the listed packages into a single shared library that will be used when building with the -linkshared option.

c-shared: (Go v1.6.0)
Requires a main package as for -buildmode=c-archive
Build the main package, plus all packages that it imports, into a single C shared library.
The only callable symbols will be those functions marked as exported.

plugin:
Requires a main package as for -buildmode=c-archive
Build the main package, plus all packages that it imports, into a single shared library that may be loaded as a run-time plugin.

exe:
Default build mode for a package named main.

pie: (Go v1.6.0)
This is like -buildmode=exe , but it builds a Position Independent Executable.


-linkshared:

Directs to Go tool to use link against shared libraries when available.
When no shared library is available for some imported package, the ordinary archive will be used instead.
The -linkshared flag may be used with
-buildmode=shared, exe, pie
As the name suggests, -linkshared is NOT used for -buildmode=archive or c-archive


Hands on:

Build Golang's std library into shared library.
$ go install -buildmode=shared std

Build the shared code:
$ go install -buildmode=shared -linkshared github.com/your/lib/code

Use the shared library:
$ go install -linkshared github.com/your/main/code  // which imports "github.com/your/lib/code"

With a SONAME, and beware to put the compiled binary into the POSIX SONAME location/version.
$ go install \
-ldflags '-extldflags -Wl,-soname,libpikachu.so.0' \
-buildmode=shared \
-linkshared \
github.com/your/lib/code

Jul 9, 2014

[C++11] constexpr with static data member array initialization

Initialize static const multidimensional array with inferred dimensions inside class definition

constexpr is very useful for compile time programming.

the difference between with constexpr and without is that

we won't have symbols while using constexpr. Everything is done

in compile time.

No symbols, no linkage, no address of the symbol being taking possible.

May 6, 2014

[C++] virtual table / RTTI during code gen and linking

Kind of forgetting this rule and what's underneath ^^

To be more clear, Mr. Honza Hubička 's nice articles mentioned:

"There are several ways to overcome redundant compilation. The most transparent one is the concept of keyed methods. If class define at least one method that is not inline, we know that there is one compilation unit defining it. Therefore all class data are located into that unit."

Provide a Virtual Method Anchor for Classes in Headers

If a class is defined in a header file and has a vtable (either it has virtual methods or it derives from classes with virtual methods), it must always have at least one out-of-line virtual method in the class. Without this, the compiler will copy the vtable and RTTI into every .o file that #includes the header, bloating .o file sizes and increasing link times.

Reference:

Jan 25, 2014

[elf] Executable Shared Libraries

Linkers and Loaders

Executable Shared Libraries

How to run a shared library on Linux

Piece of PIE

[G++][Gcc][note] compile parameters

http://goo.gl/Oz3Z49

gcc -g -c -fPIC -Wall mod1.c mod2.c mod3.c #create PIC object files
gcc -g -shared -o libfoo.so mod1.o mod2.o mod3.o #create shared library
gcc -g -fPIC -Wall mod1.c mod2.c mod3.c -shared -o libfoo.so  #create & compile share library

gcc -g -Wall -Wl,-rpath,/home/mtk/pdir -o prog prog.c libdemo.so
# –rpath : insert library path into the ELF for runtime lookup


gcc -g -Wall -o prog prog.c -Wl,--enable-new-dtags -Wl,-rpath,/home/mtk/pdir/d1 \ -L/home/mtk/pdir/d1 -lx1  #enable DT_RUNPATH in ELF

gcc -Wl,-rpath,'$ORIGIN'/lib
#to locate runtime library based on the application location



gcc -g -shared -Wl,-Bsymbolic -o libfoo.so foo.o
#To let the symbol reference to the symbol in the same library, use -Bsymbolic
# like dlopen's flag = RTLD_DEEPBIND

gcc -Wl,--export-dynamic main.c
gcc -export-dynamic main.c
#Let the library use symbols in the main program:


Runtime Library Lookup inside the ELF file

gcc -g -Wall -Wl,-rpath,/home/mtk/pdir -o prog prog.c libdemo.so
–rpath : insert library path into the ELF for runtime lookup

environment variable:
LD_RUN_PATH # used only if -rpath is NOT specified.


At run time, search for library: (precedence)
DT_RPATH //OLD in ELF
LD_LIBRARY_PATH //env variable
DT_RUNPATH   //NEW in ELF


Link Env Variable
LD_PRELOAD=libalt.so ./program
The LD_PRELOAD environment variable controls preloading on a per-process basis.

/etc/ld.so.preload #system wide basis.

LD_PRELOAD >> /etc/ld.so.preload  (precedence)
set-user-ID and set-group-ID programs ignore LD_PRELOAD.
LD_DEBUG


# grab RPATH from elf
$ readelf -d ./hello|grep RPATH
0x0000000f (RPATH) Library rpath: [/usr/local/lib/hello/B:/usr/local/lib/hello/A]


***********
The dynamic linker(runtime) looks in these places to find the library: (From book Linkers and Loaders)
If the dynamic segment contains an entry called DT_RPATH, it's a colon-separated list of directories to search for libraries. This entry is added by a command line switch or environment variable to the regular (not dynamic) linker at the time a program is linked. It's mostly used for subsystems like databases that load a collection of programs and supporting libraries into a single directory.

If there's an environment symbol LD_LIBRARY_PATH, it's treated as a colon-separated list of directories in which the linker looks for the library. This lets a developer build a new version of a library, put it in the LD_LIBRARY_PATH and use it with existing linked programs either to test the new library, or equally well to instrument the behavior of the program. (It skips this step if the program is set-uid, for security reasons.)

The linker looks in the library cache file /etc/ld.so.conf which contains a list of library names and paths. If the library name is present, it uses the corresponding path. This is the usual way that most libraries are found. (The file name at the end of the path need not be exactly the same as the library name, see the section on library versions, below.)

If all else fails, it looks in the default directory /usr/lib, and if the library's still not found, displays an error message and exits.

** LIBRARY_PATH is for linking time to search for the library, like -L option for linker!





Read the dlopen(3) man page (e.g. by typing man dlopen in a terminal on your machine):

If filename contains a slash ("/"), then it is interpreted as a (relative or absolute) pathname. Otherwise, the dynamic linker searches for the library as follows (see ld.so(8) for further details):

   o   (ELF only) If the executable file for the calling program
       contains a DT_RPATH tag, and does not contain a DT_RUNPATH tag,
       then the directories listed in the DT_RPATH tag are searched.

   o   If, at the time that the program was started, the environment
       variable LD_LIBRARY_PATH was defined to contain a colon-separated
       list of directories, then these are searched.  (As a security
       measure this variable is ignored for set-user-ID and set-group-ID
       programs.)

   o   (ELF only) If the executable file for the calling program
       contains a DT_RUNPATH tag, then the directories listed in that
       tag are searched.

   o   The cache file /etc/ld.so.cache (maintained by ldconfig(8)) is
       checked to see whether it contains an entry for filename.

   o   The directories /lib and /usr/lib are searched (in that order).


So you need to call dlopen("./libLibraryName.so", RTLD_NOW)
not just dlopen("libLibraryName.so", RTLD_NOW) which wants your plugin to be in your $LD_LIBRARY_PATH on in /usr/lib/ etc .... - or add . to your LD_LIBRARY_PATH
(don't recommend for security reasons).

Should use and display the result of dlerror when dlopen (or dlsym) fails:

  void* dlh = dlopen("./libLibraryName.so", RTLD_NOW);

  if (!dlh)

    { fprintf(stderr, "dlopen failed: %s\n", dlerror());

      exit(EXIT_FAILURE); };


Jan 24, 2014

[ELF][PIC][Note]

References and excerpt from:

http://eli.thegreenplace.net/2011/11/03/position-independent-code-pic-in-shared-libraries/

http://eli.thegreenplace.net/2011/11/11/position-independent-code-pic-in-shared-libraries-on-x64/

http://eli.thegreenplace.net/2011/08/25/load-time-relocation-of-shared-libraries/

http://eli.thegreenplace.net/2012/01/03/understanding-the-x64-code-models/

http://eli.thegreenplace.net/2012/08/13/how-statically-linked-programs-run-on-linux/



ecx : Serves as a base pointer to GOT.  e.g : 0x1ff4

a value is taken from [ecx - 0x10], which is a GOT entry, and placed into eax.

The address of myglob : ecx - 0x10 = 0x1fe4 = eax

eax:  address of myglob. eax address : 0x1fe4.

library that uses myglob is pointed to here.

eax type: R_386_GLOB_DAT : put the actual value of the symbol
  (i.e. its address) into that offset".

eax : the value of myglob is placed into eax.

--------
Function:

lazy binding optimization:
When a shared library refers to some function, the real address of that function
  is not known until load time.

Lazy binding scheme is attained by adding yet another level of indirection – the PLT.

Procedure Linkage Table (PLT):
PLT is part of the executable text section, consisting of a set of entries.
(one for each external function the shared library calls)
Each PLT entry is a short chunk of executable code.

Instead of calling the function directly, the code calls an entry in the PLT,
which then takes care to call the actual function.

This arrangement is sometimes called a "trampoline".

Each PLT entry also has a corresponding entry in the GOT which contains the actual
offset to the function, but only when the dynamic loader resolves it.

When the shared library is first loaded, the function calls have not been resolved yet:
PLT[n] => GOT[n] , GOT[n] Has the function's address


PLT[0] is different from other entry. It's a call to a resolver routine,
which is located in the dynamic loader itself.
This routine resolves the actual address of the function.
-------
Flow:

First time call:
Text call function func: -> PLT[n] -> GOT[n] -> back to next address of PLT[n]'s
call to GOT[n], which is to prepare resolver -> PLT[0] , call dynamic loader,
prepare function address and place into GOT[n] -> calls the function.

Next time call:
Text call function func: -> PLT[n] -> GOT[n] ->  calls the function.
------

Lazy symbol resolution performed by the dynamic loader can be configured with
some environment variables:

LD_BIND_NOW : always perform the resolution for all symbols at start-up time.
  With GDB. You’ll see that the GOT entry for ml_util_func contains its real address even before the first call to the function.

LD_BIND_NOT : not to update the GOT entry at all.

More info:
man ld.so

------
The costs of PIC:
1. extra indirection required for all external references to data and code in PIC.
2. increased register usage required to implement PIC.
  it makes sense for the compiler to generate code that keeps its address in a register
  (usually ebx).

------
X64:
RIP-relative addressing:
New "RIP-relative addressing mode",the default for all 64-bit mov instructions that reference memory
(it’s used for other instructions as well, such as lea).
A quote from the "Intel Architecture Manual vol 2a":
A new addressing form, RIP-relative (relative instruction-pointer) addressing, is implemented in 64-bit mode.
An effective address is formed by adding displacement to the 64-bit RIP of the next instruction.

The displacement used in RIP-relative mode is 32 bits in size.
Since it should be useful for both positive and negative offsets,
roughly +/- 2GB is the maximal offset from RIP supported by this addressing mode.

Jan 8, 2014

[C++11][NOTE] std::vector performance regression when enabling C++11

std::vector performance regression when enabling C++11

GCC/G++ LinkTimeOptimization

gcc -flto -o f f1.o f2.o

Link-time optimization does not work well with generation of debugging information.
Combining -flto with -g is currently experimental and expected to produce wrong results.